{"id":"7ba32980-1698-48f7-857a-ee02a4fa3b40","entityType":"agent","slug":"clawhub-zw008-inference-aiops","name":"inference-aiops","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-inference-aiops","canonicalPath":"/agent/clawhub-zw008-inference-aiops","generatedAt":"2026-10-10T10:44:06.314Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":null},"description":"Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens. Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster. Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray). Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:inference-aiops","sourceUrl":"https://clawhub.ai/zw008/inference-aiops","homepage":"https://clawhub.ai/zw008/skills/inference-aiops","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/inference-aiops","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/inference-aiops","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":64,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"inference-aiops technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":null},"stars":null,"forks":null,"downloads":1596,"packageName":null,"latestVersion":"0.10.4","tractionLabel":"1.6K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T07:16:11.212Z","lastCrawledAt":"2026-10-10T07:16:11.212Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T07:16:11.212Z","lastVerifiedAt":null,"highlights":[{"version":"0.10.4","createdAt":"2026-09-16T23:25:50.857Z","changelog":"- Updated documentation in references/agent-guardrails.md. - Removed the skill-card.md file. - No functional changes to skill logic; documentation and metadata only.","fileCount":7,"zipByteSize":19850},{"version":"0.10.3","createdAt":"2026-09-15T06:01:11.929Z","changelog":"- Removed the file: skill-card.md - No changes to functionality or behavior - Documentation or metadata cleanup only","fileCount":7,"zipByteSize":19596},{"version":"0.10.2","createdAt":"2026-09-12T14:25:06.436Z","changelog":"- Removed the file skill-card.md from the repository. - No functional code changes; documentation (SKILL.md) remains largely unchanged. - Housekeeping update to prune redundant project files.","fileCount":7,"zipByteSize":19570},{"version":"0.10.1","createdAt":"2026-09-12T10:08:58.527Z","changelog":"inference-aiops 0.10.1 - Documentation revised in SKILL.md; content updated for clarity, no functional/tool changes. - skill-card.md file removed. - No changes to user-facing features, tools, or APIs.","fileCount":7,"zipByteSize":19747},{"version":"0.10.0","createdAt":"2026-09-12T00:57:30.083Z","changelog":"- Updated dependency requirements: now requests either \"inference-aiops\" or \"uvx\" as valid CLI binaries. - Updated metadata: primary required environment variable and config attributes adjusted for flexibility, matching launcher requirements. - Removed the file skill-card.md from the repository. - Example \"Quick Install\" snippet updated for clarity and correctness.","fileCount":7,"zipByteSize":19381},{"version":"0.9.0","createdAt":"2026-08-10T06:51:16.474Z","changelog":"- Removed the file: skill-card.md - No other functional changes introduced in this version.","fileCount":7,"zipByteSize":19425},{"version":"0.8.0","createdAt":"2026-08-03T05:53:02.421Z","changelog":"- Removed redundant file: skill-card.md - No user-facing functionality changed - Documentation and core features remain unchanged - Maintenance release to simplify project files","fileCount":7,"zipByteSize":19545},{"version":"0.7.0","createdAt":"2026-08-02T09:39:46.635Z","changelog":"- Removed the skill documentation file skill-card.md. - No changes to functionality or APIs; this is a documentation/pruning-only update. - All existing features and governance behavior remain unchanged.","fileCount":7,"zipByteSize":19678}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:inference-aiops","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:inference-aiops` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/inference-aiops before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T10:44:06.309Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":null},"readme":"Skill: inference-aiops\n\nOwner: zw008\n\nSummary: Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens. Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster. Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray). Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).\n\nTags: agent-skills:0.1.0, ai-ops:0.1.0, latest:0.10.4, mcp:0.1.0, ray:0.1.0, vllm:0.1.0\n\nVersion history:\n\nv0.10.4 | 2026-09-16T23:25:50.857Z | auto\n\n- Updated documentation in references/agent-guardrails.md.\n- Removed the skill-card.md file.\n- No functional changes to skill logic; documentation and metadata only.\n\nv0.10.3 | 2026-09-15T06:01:11.929Z | auto\n\n- Removed the file: skill-card.md\n- No changes to functionality or behavior\n- Documentation or metadata cleanup only\n\nv0.10.2 | 2026-09-12T14:25:06.436Z | auto\n\n- Removed the file skill-card.md from the repository.\n- No functional code changes; documentation (SKILL.md) remains largely unchanged.\n- Housekeeping update to prune redundant project files.\n\nv0.10.1 | 2026-09-12T10:08:58.527Z | auto\n\ninference-aiops 0.10.1\n\n- Documentation revised in SKILL.md; content updated for clarity, no functional/tool changes.\n- skill-card.md file removed.\n- No changes to user-facing features, tools, or APIs.\n\nv0.10.0 | 2026-09-12T00:57:30.083Z | auto\n\n- Updated dependency requirements: now requests either \"inference-aiops\" or \"uvx\" as valid CLI binaries.\n- Updated metadata: primary required environment variable and config attributes adjusted for flexibility, matching launcher requirements.\n- Removed the file skill-card.md from the repository.\n- Example \"Quick Install\" snippet updated for clarity and correctness.\n\nv0.9.0 | 2026-08-10T06:51:16.474Z | auto\n\n- Removed the file: skill-card.md\n- No other functional changes introduced in this version.\n\nv0.8.0 | 2026-08-03T05:53:02.421Z | auto\n\n- Removed redundant file: skill-card.md\n- No user-facing functionality changed\n- Documentation and core features remain unchanged\n- Maintenance release to simplify project files\n\nv0.7.0 | 2026-08-02T09:39:46.635Z | auto\n\n- Removed the skill documentation file skill-card.md.\n- No changes to functionality or APIs; this is a documentation/pruning-only update.\n- All existing features and governance behavior remain unchanged.\n\nv0.6.0 | 2026-07-21T09:41:16.487Z | auto\n\n- Governance audit log entries now include descriptive risk-tier labels instead of authorization gating.\n- `@governed_tool` decorator updated to record operation budget/runaway status and risk-tiers, clarifying that it records rather than authorizes actions.\n- Documentation updated for governance harness behavior and audit fields.\n- `skill-card.md` has been removed.\n\nv0.5.0 | 2026-07-20T11:15:30.270Z | auto\n\n**Expanded toolset and lifecycle enhancements in inference-aiops 0.5.0:**\n\n- Tool count increased from 37 to 39, with expanded read and write operations (now 23 read, 16 write).\n- Added new tools in the Metrics & RCA, Ray Serve, and deploy lifecycle categories.\n- Introduced model_sleep and additional reversible write actions, improving operational safety.\n- Updated documentation and references to reflect new tool counts and usage patterns.\n- Removed obsolete skill-card.md file.\n\nv0.4.0 | 2026-07-19T03:51:35.518Z | auto\n\nVersion 0.4.0 of inference-aiops\n\n- Added new \"Agent Guardrails\" reference documentation.\n- Documentation improvements: clarified skill purpose, cleaned up tags and summary, and improved organization in SKILL.md.\n- Expanded toolset description: now explicitly mentions all 37 MCP tools.\n- Updated compatibility and validation disclaimer for more accurate status.\n- Removed outdated file (skill-card.md).\n\nv0.3.0 | 2026-07-17T05:55:27.966Z | auto\n\n**inference-aiops 0.3.0 adds support for SGLang and TGI inference engines alongside vLLM/Ray.**\n\n- Adds engine-agnostic observability and root-cause tools for SGLang and TGI (health, inventory, request metrics, queue depth, latency diagnostics)\n- CLI config wizard now supports vLLM, SGLang, and TGI targets\n- New engine-agnostic tool group: diagnose_engine_latency, engine health, inventory, and queue\n- Tool count increased (now 35 tools: 21 read, 14 write)\n- Updated documentation and removed obsolete skill-card.md\n- vLLM-only scale/drain/deploy actions now clearly teach-and-refuse on SGLang/TGI targets\n\nv0.2.0 | 2026-07-13T13:08:40.420Z | auto\n\ninference-aiops 0.2.0\n\n- Removed the skill-card.md file (now only SKILL.md remains as the main documentation).\n- No changes to features or functionality.\n- Documentation and metadata cleanup: the SKILL.md is now the single source of detailed info.\n\nv0.1.0 | 2026-07-12T06:59:51.276Z | auto\n\nInitial preview release of inference-aiops: a governed toolkit for GPU inference cluster operations (vLLM & Ray Serve).\n\n- Provides 30 tools for monitoring, diagnosing, and controlling vLLM and Ray Serve clusters.\n- Includes full governance harness: audit log, policy engine, undo, budget caps, risk tiers.\n- Supports metrics, root-cause analysis, scaling, deployment lifecycle, LoRA/model management, GPU/job/routing control, and cost per token.\n- All sensitive credentials stored encrypted; write operations require confirmation and support dry-run/undo.\n- No external server dependencies; parses vLLM /metrics directly and operates standalone.\n- PREVIEW: mock-validated only, not yet production-tested on live clusters.\n\nArchive index:\n\nArchive v0.10.4: 7 files, 19850 bytes\n\nFiles: references/agent-guardrails.md (7485b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (2842b), SKILL.md (17210b), _meta.json (135b)\n\nFile v0.10.4:SKILL.md\n\n---\nname: inference-aiops\nslug: inference-aiops\ndisplayName: \"Inference AIops\"\nsummary: \"Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Inference-AIops\ntags: [aiops, mcp, governance, inference]\ndescription: >\n  Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens.\n  Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster.\n  Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray).\n  Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).\ninstaller:\n  kind: uv\n  package: inference-aiops\nargument-hint: \"[deployment/model name or describe your inference-cluster task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"inference-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"INFERENCE_AIOPS_CONFIG\",\"INFERENCE_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Inference-AIops\",\"emoji\":\"🚀\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.inference-aiops/ (relocatable via INFERENCE_AIOPS_HOME).\n  Auth: a bearer token is OPTIONAL — many vLLM / Ray stacks run open. When the API requires one it is stored ENCRYPTED in ~/.inference-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'inference-aiops init' to onboard, or 'inference-aiops secret set <target>' to add one. The store is unlocked by a master password from INFERENCE_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var INFERENCE_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback (migrate with 'inference-aiops secret migrate'). The token is sent as an Authorization: Bearer header at request time and held only in memory; it is never logged or echoed.\n  State-changing operations require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + risk-tier label — it records, not authorizes). The fragile prod ops — scale_replicas_down, scale_to_zero, drain_replica, lora_unload, model_sleep, replica_restart, model_undeploy, deployment_redeploy — are high-risk with a dry_run preview; reversible writes (scale, autoscale-config, routing, sleep, LoRA load) record an undo descriptor.\n  Engines: vLLM (with its Ray Serve control plane), SGLang, and TGI. SGLang/TGI are single-process servers with engine-agnostic observability (health, running-model inventory, request metrics, queue depth, latency RCA); Ray-shaped scale/drain writes are vLLM-only and raise a teaching error on a SGLang/TGI target.\n  Metrics: each engine's Prometheus /metrics endpoint is parsed directly — no Prometheus server is required.\n  Webhooks: none — no outbound calls beyond the configured Ray dashboard and vLLM services.\n  SSL: verify_ssl defaults to true; disable only for self-signed lab certificates.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked vLLM/Ray responses; unverified against multi-GPU tensor/pipeline-parallel deployments, real GPU thermal/throttle telemetry, and multi-node drain (see docs/VERIFICATION.md).\n---\n\n# Inference AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor.** Product and trademark names belong to their owners. Source at [github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops) under the MIT license.\n\nGoverned GPU-inference operations for **vLLM** (OpenAI API + Prometheus `/metrics`) and **Ray Serve / Ray Jobs** (Ray dashboard), plus the single-process serving engines **SGLang** and **TGI** — **39 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.inference-aiops/`, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels on every audit row. The flagship `diagnose_latency_spike` folds queue depth + KV-cache pressure + prefix-cache locality into a ranked cause and the specific knob to turn; the engine-agnostic `diagnose_engine_latency` does the same across whatever signals SGLang/TGI expose. Each engine's Prometheus `/metrics` is parsed directly — **no Prometheus server required**.\n\n> **Standalone**: the governance harness is bundled in the package (`inference_aiops.governance`) — no external skill-family dependency. A bearer token is **optional** (many stacks run open).\n\n## What This Skill Does\n\n| Group | Tools | Count | Read or Write |\n|-------|-------|:-----:|:-------------:|\n| **Metrics & RCA** (vLLM) | request metrics, queue depth, KV-cache stats, diagnose latency spike, diagnose low utilisation | 5 | 5 read |\n| **Engine-agnostic** (vLLM/SGLang/TGI) | engine health, engine inventory, engine request metrics, engine queue depth, diagnose engine latency | 5 | 5 read |\n| **Ray Serve (read)** | deployment list, deployment status, replica list, autoscale config get | 4 | 4 read |\n| **Ray Serve (write)** | scale up (med), scale down (high), scale-to-zero (high), autoscale config update (med), drain replica (high) | 5 | 5 write |\n| **Models / vLLM** | model list, model info, LoRA load (med), LoRA unload (high), base hot-swap (high) | 5 | 2 read / 3 write |\n| **Ray cluster / jobs / GPU** | cluster resources, dashboard status, job list, GPU utilisation, job cancel (med), replica restart (high) | 6 | 4 read / 2 write |\n| **Deploy lifecycle** | deploy (med), undeploy (high), redeploy (high), routing policy update (med) | 4 | 4 write |\n| **Cost** | cost per token | 1 | 1 read |\n\n**23 read, 16 write**, plus `undo_list` / `undo_apply` — **39 MCP tools** in total. The high-risk writes support `dry_run` + double-confirm; reversible writes record an undo descriptor. The engine-agnostic reads cover any engine; the Ray Serve / cluster / deploy write groups are vLLM-only and teach-and-refuse on a SGLang/TGI target (single-process engines have no Ray control plane).\n\n## Quick Install\n\n```bash\nuv tool install inference-aiops\ninference-aiops init       # interactive wizard: engine (vllm/sglang/tgi) + host + port + scheme (token optional)\ninference-aiops doctor     # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/inference-aiops\nopenclaw skills info inference-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Triage a cluster (`overview`): Serve deployments, total replicas, queue backpressure\n- Diagnose slow inference (`metrics diagnose` / `diagnose_latency_spike`): rank the cause (queue depth vs KV-cache preemption vs prefix-cache locality) and get the knob to turn\n- Find idle GPUs and over-provisioned replicas (`diagnose_low_utilization`)\n- Scale a Ray Serve deployment up/down, **scale-to-zero** to stop cost bleed, or update autoscale bounds\n- **Drain** a replica gracefully before a node reboot (finishes in-flight requests)\n- Load/unload a **LoRA** adapter; **hot-swap** a base model (Sleep-Mode swap, captures the prior model)\n- Inspect GPU utilisation per node, list/cancel Ray jobs, restart a stuck replica\n- Compute **cost per million tokens** from throughput × GPU $/hr\n- Observe an **SGLang** or **TGI** server (`engine_health`, `engine_inventory`, `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) — single-process engines with no Ray control plane\n\n**Do NOT use for** non-inference infrastructure (hypervisors, storage appliances, backup products, general container workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools. This skill is scoped to GPU inference serving (vLLM + Ray).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| vLLM / Ray Serve inference: latency RCA, autoscale, drain, LoRA, cost/token | **inference-aiops** (this skill) |\n| SGLang / TGI serving: health, running-model inventory, request metrics, queue depth, latency RCA | **inference-aiops** (this skill — engine-agnostic reads) |\n| Any non-inference infrastructure (hypervisor, storage, backup, general clusters, network, OT) | the appropriate **other AIops-tools** line |\n\n## Common Workflows\n\n### 1. \"Inference got slow this afternoon\" (flagship RCA → the right knob)\n\n1. `inference-aiops doctor` → confirm the vLLM endpoint and Ray dashboard are actually reachable before blaming the model\n2. `inference-aiops overview` → Serve deployments, total replicas, and whether queue backpressure is cluster-wide or one deployment\n3. `inference-aiops metrics diagnose` (MCP: `diagnose_latency_spike`) → a **ranked** cause with the measured numbers: is `waiting` queue depth high (backpressure)? Are there KV-cache **preemptions** (`kv_cache_stats`)? Has the **prefix-cache hit rate** dropped (routing lost locality)?\n4. Turn the knob the RCA names, not a guess:\n   - backpressure → `inference-aiops serve scale <app> <deployment> --replicas N` (`scale_replicas_up`, reversible, prior count captured)\n   - KV-cache preemption → `autoscale_config_update` to lower the concurrent-request cap (reversible, prior config captured)\n   - lost locality → `routing_policy_update` to prefix-aware / session-affinity (reversible)\n5. Re-check `inference-aiops metrics requests` (TTFT / TPOT / e2e) and `inference-aiops metrics queue` to confirm the p99 actually moved\n6. **Failure branch**: if the fix makes it worse, `inference-aiops undo list` → `inference-aiops undo apply <id>` restores the exact prior replica count / autoscale config / routing policy. If `diagnose_latency_spike` reports no clear cause, the bottleneck is likely upstream of serving — check `gpu_utilization` for a throttling or shared-GPU problem before scaling anything.\n\n### 2. Off-peak cost save: scale a deployment down to zero and bring it back\n\n1. `inference-aiops metrics requests` → confirm traffic really is idle, not just briefly quiet\n2. `diagnose_low_utilization` → the deployments actually burning GPU for nothing, with the measured utilisation\n3. `cost_per_token` → quantify the bleed ($/1M tokens at the current throughput) so the change is justifiable in the audit trail\n4. (optional) `export INFERENCE_AUDIT_APPROVED_BY=you INFERENCE_AUDIT_RATIONALE=\"off-peak cost save\"` → annotates the audit row with who/why; recorded when set, never required\n5. `inference-aiops serve scale-to-zero <app> <deployment> --dry-run`, then re-run without `--dry-run` → **high** risk, double confirmation. `scale_to_zero` stops the bleed but **strands ingress** — requests will queue or fail until replicas return\n6. To restore: `inference-aiops undo apply <id>` (replays the captured prior replica count) or `inference-aiops serve scale <app> <deployment> --replicas N`\n7. **Failure branch**: if traffic arrives while at zero, restore immediately via undo — do not wait for autoscale, since `scale_to_zero` may have been applied outside the autoscaler's floor. If the restore fails, `serve status` will show the deployment unhealthy; `deployment_redeploy` is the last resort (high risk, disruptive).\n\n### 3. Drain a replica before a node reboot\n\n1. `inference-aiops serve list` / `replica_list` → identify the replicas pinned to the node you are about to reboot\n2. `queue_depth` → confirm the remaining replicas can absorb the load; if not, `scale_replicas_up` **first** so draining does not cause a brownout\n3. `drain_replica <app> <deployment> <replica_id> --dry-run`, then confirm → **high** risk; the drain finishes in-flight requests before removing the replica\n4. Watch `replica_list` until the replica is gone and `request_metrics` shows no error spike, then reboot the node\n5. **Failure branch**: if the drain hangs on a long-running request, `replica_restart` forcibly cycles it — that **drops** in-flight requests, so only reach for it once you accept the loss. Multi-node drain has not been verified against a live cluster (see `docs/VERIFICATION.md`).\n\n### 4. Free GPU memory between bursts with Sleep Mode, then resume\n\n1. `model_is_sleeping` → is the engine already suspended? `null` means the engine did not report it — that is UNKNOWN, not awake, so resolve it before writing\n2. `request_metrics` / `queue_depth` → confirm the engine is actually idle; sleeping a busy engine drops live traffic\n3. `model_sleep --dry-run`, then confirm → **high** risk. Level 1 offloads the weights to CPU RAM and wakes fast; level 2 discards them, so waking reloads from disk. The undo descriptor is recorded **only** if the engine was observed awake first — an already-sleeping engine records none, so an undo can never wake something this call did not suspend\n4. Verify: `model_is_sleeping` reports true, and GPU memory has been released (`gpu_utilization`)\n5. Resume with `model_wake` (medium risk), or `inference-aiops undo apply <id>` to replay the recorded inverse. `model_wake` itself records **no** undo: vLLM reports whether the engine sleeps but never at which level, and guessing between level 1 and level 2 would be inventing a prior state\n6. **Failure branch**: if any of the three tools reports that the route does not exist, the server was **not** started with `VLLM_SERVER_DEV_MODE=1`. That is a server start-up flag, not a fault in the tool and not a stale id — restart vLLM with the flag, or leave Sleep Mode off if this is a production deployment that should not expose it.\n\n> vLLM has **no** in-place base-model swap. Sleep Mode suspends and resumes the *same* model; serving a different base model means restarting vLLM with a different `--model`. For adapter-level changes use `lora_load` (reversible) and `lora_unload` (high).\n\n## Governance & Safety\n\nThe skill delivers reads and writes and records them; it does **not** decide\nwhether a write is permitted. That is your agent's judgement, or the permission\nof the environment you connect it with (a network path that only reaches the\nread/metrics endpoints, a Ray dashboard without its job-submission API — writes\nthen fail at the server). There is no read-only switch, policy file, or approval\ngate.\n\n- **Audit is the guarantee, and it is not bypassable.** Every operation — MCP and CLI alike — is logged to `~/.inference-aiops/audit.db` (relocatable via `INFERENCE_AIOPS_HOME`): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.\n- `INFERENCE_AUDIT_APPROVED_BY` / `INFERENCE_AUDIT_RATIONALE` are optional annotations recorded on the audit row (who/why); they are never required and never block.\n- **Runaway guard** — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker.\n- The fragile prod writes support `--dry-run` / `dry_run=True` and double confirmation at the CLI.\n- Reversible writes (scale, autoscale-config, routing, hot-swap, LoRA load) capture before-state and record an inverse descriptor.\n\n## References\n\n- `references/capabilities.md` — full tool → backend → endpoint → returns reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, optional token, and connectivity\n\nFile v0.10.4:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"inference-aiops\",\n  \"version\": \"0.10.4\",\n  \"publishedAt\": 1789601150857\n}\n\nFile v0.10.4:references/agent-guardrails.md\n\n# Agent guardrails — running inference-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the environment's. The tool\ndoes not gate it — there is no read-only switch and no approval prompt to\nconfigure. The two right places to control read vs write:\n\n- **The environment you connect it to.** Restrict the network path so the tool\n  can only reach the read/metrics endpoints, or run the Ray dashboard without its\n  job-submission API. A write then fails at the server, which is the only place\n  the permission actually lives — no skill-side flag can be argued around by a\n  model, but a blocked endpoint cannot.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the engine or Ray dashboard did not return comes back as `null`, never as `\"\"`. An absent job `entrypoint`, a model's `parent` adapter, a replica `state`, or a server-info `version` is distinguishable from an empty one. |\n| \"Tell me if the output was cut off\" | `ray_job_list` returns `{\"jobs\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`. Truncation is measured against the full fetch, not guessed from a length coincidence. |\n| \"Say when a metric isn't available\" | Signals the engine does not expose come back as `null` rather than `0`. SGLang and TGI expose fewer metrics than vLLM; `diagnose_engine_latency` skips a signal it cannot read instead of fabricating it, and `signalsChecked` shows exactly what it looked at. |\n| \"Don't suggest scaling on an engine that can't scale\" | Multi-replica scale / drain / autoscale are Ray Serve control-plane actions. On a single-process engine (SGLang, TGI) those tools raise `EngineCapabilityError` with an explanation, rather than issuing a call that could never succeed. |\n| \"Confirm before anything disruptive\" | Every traffic-affecting operation (`model_undeploy`, `deployment_redeploy`, `scale_to_zero`, `scale_replicas_down`, `drain_replica`, `replica_restart`, `lora_unload`, `model_sleep`) takes `dry_run=True` for a preview and is `risk=high`. ⚠️ **The double confirmation is a CLI feature, and only `scale_to_zero` has a CLI command** — every other one is reachable only over MCP, where nothing prompts. Keep your own confirmation for those. |\n| \"Log what you did\" | Every call is audited to `~/.inference-aiops/audit.db` regardless of what the model says it did. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a GPU inference cluster through the inference-aiops MCP tools\n(vLLM / SGLang / TGI serving engines, plus a Ray Serve control plane).\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a tool.\n  Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"healthy\".\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null metric means the engine does not expose that signal. Report it as \"not\n  available\" — never substitute 0, and never compare a null against a threshold.\n- Report values exactly as returned. Do not normalise or prettify model ids,\n  deployment names, replica states, or Ray job statuses.\n- When diagnose_engine_latency or diagnose_latency_spike returns probableCauses,\n  work through them in the order given and cite the measured number in each\n  cause's \"signal\" — do not substitute your own theory of the bottleneck.\n\n- Only `scale_to_zero` has a CLI command; every other traffic-affecting tool is MCP-only\n  and nothing will ask you to confirm it. Call it with `dry_run=True` first, show the\n  operator what would change, and wait for an explicit go-ahead.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a latency, throughput, or capacity problem unless a tool result\n  supports it. High GPU utilisation is not by itself a fault.\n- Do not confuse the identifier kinds: a Ray *application* name, a *deployment*\n  name within it, a *replica* id, a Ray *job* id (raysubmit_…), and a served\n  *model* id are four different things. Never pass one where another is expected.\n- cost_per_token is arithmetic over a price you supply, not a billing figure.\n  Present it as an estimate with its inputs.\n```\n\n## Recommended setup for a local model\n\nStart with a path that *cannot* write — restrict the network route to the\nread/metrics endpoints, or expose the Ray dashboard without its job-submission\nAPI — verify, and widen access only when you trust the setup. The\ntraffic-affecting operations here (`scale_to_zero`, `drain_replica`,\n`model_undeploy`) strand or drop live requests and are cheap to invoke:\n\n```bash\ninference-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport INFERENCE_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport INFERENCE_AUDIT_RATIONALE=\"scaling llm-app down for the maintenance window\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the `diagnose_*` tools —\n  `diagnose_engine_latency`, `diagnose_latency_spike`, `diagnose_low_utilization`\n  do the multi-signal correlation inside one call, so the model does not have to\n  chain reads and keep deployment/replica ids straight.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions and use `limit` deliberately rather than dumping a cluster's whole\n  job history.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.4:references/capabilities.md\n\n# inference-aiops capabilities\n\n> 39 MCP tools (23 read, 16 write, 2 undo). Serving engines:\n> **vLLM** (OpenAI API + Prometheus `/metrics`, default 8000) with its **Ray**\n> dashboard control plane (Serve + Jobs, default 8265), plus the single-process\n> engines **SGLang** (OpenAI API + `/get_server_info` + Prometheus `/metrics`,\n> default 30000) and **TGI** (`/info` + Prometheus `/metrics`, default 8080).\n> Endpoints modelled against those APIs; need live verification.\n\n## Metrics & RCA — vLLM (read, 5)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `request_metrics` | vLLM | `GET /metrics` | TTFT, TPOT, e2e latency (avg/p50/p90/p99), prompt/generation token totals, request counts |\n| `queue_depth` | vLLM | `GET /metrics` | running vs waiting requests (backpressure), scheduler state |\n| `kv_cache_stats` | vLLM | `GET /metrics` | KV-cache utilisation %, prefix-cache hit rate, preemption count |\n| `diagnose_latency_spike` | vLLM | `GET /metrics` (fold) | **ranked cause** (queue backpressure / KV-cache preemption / prefix-cache locality) + the specific knob to turn |\n| `diagnose_low_utilization` | vLLM | `GET /metrics` (fold) | idle-GPU / over-provisioned / routing-stranded diagnosis + what to scale down |\n\n## Engine-agnostic — vLLM / SGLang / TGI (read, 5)\n\nWork against **any** supported engine, reading each engine's own paths and metric\nnames (vLLM `vllm:*`, SGLang `sglang:*`, TGI `tgi_*`). A signal an engine does not\nexpose (e.g. TGI has no TTFT or KV-cache metric) degrades to `null` rather than\nbeing guessed.\n\n| Tool | Endpoint(s) | Returns |\n|------|-------------|---------|\n| `engine_health` | `GET /health` | engine liveness (`healthy` bool) + engine label |\n| `engine_inventory` | `GET /v1/models` (vLLM/SGLang) or `/info` (TGI); `/get_server_info` (SGLang) | running-model id(s) + best-effort server info (model, version, max concurrency) |\n| `engine_request_metrics` | `GET /metrics` | TTFT / TPOT / e2e latency + generation-token totals, per engine's exposition (null where unexposed) |\n| `engine_queue_depth` | `GET /metrics` | running vs waiting requests + backpressure flag (SGLang `num_queue_reqs`, TGI `tgi_queue_size`) |\n| `diagnose_engine_latency` | `GET /metrics` (fold) | **ranked cause** across the signals the engine exposes (queue backpressure / KV-token-cache pressure / cache locality) + the knob to turn |\n\n## Ray Serve — read (4)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `serve_deployment_list` | Ray | `GET /api/serve/applications/` | all Serve deployments: status, replica count, target |\n| `deployment_status` | Ray | `GET /api/serve/applications/` | one deployment's status + current/target replica count |\n| `replica_list` | Ray | `GET /api/serve/applications/` | per-replica id, state, node |\n| `autoscale_config_get` | Ray | `GET /api/serve/applications/` | min/max replicas, target ongoing requests |\n\n## Ray Serve — write (5)\n\n| Tool | Risk | Backend | Endpoint | Undo / safety |\n|------|------|---------|----------|---------------|\n| `scale_replicas_up` | med | Ray | `PUT /api/serve/applications/{app}/deployments/{dep}` | reversible (records prior count) |\n| `scale_replicas_down` | **high** | Ray | `PUT …/deployments/{dep}` | dry-run; captures prior count → undo |\n| `scale_to_zero` | **high** | Ray | `PUT …/deployments/{dep}` | dry-run; stops cost bleed but **strands ingress**; captures prior count → undo |\n| `autoscale_config_update` | med | Ray | `PUT …/deployments/{dep}/autoscale` | reversible (records prior bounds) |\n| `drain_replica` | **high** | Ray | `POST …/deployments/{dep}` (drain) | dry-run; graceful — finishes in-flight requests; no undo |\n\n## Models / vLLM (5)\n\n| Tool | R/W (risk) | Backend | Endpoint | Notes |\n|------|-----------|---------|----------|-------|\n| `model_list` | read | vLLM | `GET /v1/models` | served model ids |\n| `model_info` | read | vLLM | `GET /v1/models` | one model's detail (normalised) |\n| `model_is_sleeping` | read | vLLM | `GET /is_sleeping` | dev-mode only; `null` = engine did not say (UNKNOWN, not awake) |\n| `lora_load` | write (med) | vLLM | `POST /v1/load_lora_adapter` | reversible (undo unloads it) |\n| `lora_unload` | write (**high**) | vLLM | `POST /v1/unload_lora_adapter` | dry-run |\n| `model_sleep` | write (**high**) | vLLM | `POST /sleep?level=N` | dev-mode only; dry-run; level 1 offloads weights to CPU RAM, 2 discards them; captures `wasSleeping` → undo wakes only what it suspended |\n| `model_wake` | write (med) | vLLM | `POST /wake_up` | dev-mode only; dry-run; records **no** undo — vLLM never reports the prior sleep *level*, so re-sleeping would be a guess |\n\n> **Sleep Mode needs `VLLM_SERVER_DEV_MODE=1`.** vLLM registers `/sleep`, `/wake_up`\n> and `/is_sleeping` only under that flag, so on a normal production server these\n> three tools report that the route does not exist and why — a 404 here means the\n> server was not started in dev mode, not that an id was stale. Sleep Mode suspends\n> the **same** model; vLLM has no in-place base-model swap (that needs a restart\n> with a different `--model`).\n\n## Ray cluster / jobs / GPU (6)\n\n| Tool | R/W (risk) | Backend | Endpoint | Returns / notes |\n|------|-----------|---------|----------|-----------------|\n| `ray_cluster_resources` | read | Ray | `GET /api/cluster_status` | CPU/GPU total vs allocated |\n| `ray_dashboard_status` | read | Ray | Ray dashboard | dashboard reachability/version |\n| `ray_job_list` | read | Ray | `GET /api/jobs/` | submitted jobs + status |\n| `gpu_utilization` | read | Ray | `GET /api/nodes` | per-node GPU count, utilisation %, memory (best-effort) |\n| `ray_job_cancel` | write (med) | Ray | `POST /api/jobs/{id}/stop` | cancel a running job |\n| `replica_restart` | write (**high**) | Ray | `GET/PUT /api/serve/applications/` | dry-run; restart a stuck replica |\n\n## Deploy lifecycle (4, write)\n\n| Tool | Risk | Backend | Endpoint | Undo / safety |\n|------|------|---------|----------|---------------|\n| `model_deploy` | med | Ray | `PUT /api/serve/applications/` | deploy an application |\n| `model_undeploy` | **high** | Ray | `DELETE/PUT /api/serve/applications/` | dry-run |\n| `deployment_redeploy` | **high** | Ray | `PUT /api/serve/applications/` | dry-run |\n| `routing_policy_update` | med | Ray | `PUT /api/serve/applications/` | reversible; prefix-aware / session-affinity routing to fix cache locality |\n\n## Cost (read, 1)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `cost_per_token` | vLLM | `GET /metrics` + GPU $/hr | deterministic $/1M tokens from measured throughput × GPU hourly rate |\n\n## SGLang / TGI writes (control-plane teaching error)\n\nSGLang and TGI are **single-process servers** with no Ray Serve control plane, so\nthe Ray-shaped write groups above (scale / drain / autoscale / deploy / redeploy /\nrouting / job-cancel / replica-restart) do not apply to them. Attempting one\nagainst a SGLang/TGI target raises `EngineCapabilityError` with a teaching message\npointing at a real horizontal-scale layer (Ray Serve / Kubernetes / a load\nbalancer). Their supported surface is the **engine-agnostic read** group above.\n\n## Out of scope (by design)\n\n- Cluster **provisioning** (spinning up GPU nodes, driver/CUDA install)\n- vLLM/Ray **install or version upgrades**\n- Multi-node **drain / reboot orchestration** (single-replica drain only, unverified at multi-node scale)\n- Non-inference infrastructure (use the appropriate other AIops-tools line)\n\nWant one of these? Open an issue or PR — feedback and contributions welcome.\n\nFile v0.10.4:references/cli-reference.md\n\n# inference-aiops CLI reference\n\n> Serving engines: vLLM (OpenAI API + Prometheus `/metrics`,\n> default 8000) with its Ray dashboard control plane (Serve + Jobs, default 8265),\n> plus single-process SGLang (default 30000) and TGI (default 8080); endpoints\n> need live verification.\n>\n> The CLI is a convenience subset. The full 35-tool surface — including the\n> engine-agnostic reads (`engine_health`, `engine_inventory`,\n> `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) that\n> cover SGLang/TGI — is via the MCP server (`inference-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\ninference-aiops init                      # interactive wizard: engine (vllm/sglang/tgi) + host + port\ninference-aiops doctor [--skip-auth]      # config + secret store + connectivity — vLLM: Ray + vLLM; SGLang/TGI: engine health + inventory\ninference-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.inference-aiops/secrets.enc — only if a token is used)\n\n```bash\ninference-aiops secret set <target> [--value <token>]  # store a bearer token (hidden prompt if no --value)\ninference-aiops secret list                            # names only — values never shown\ninference-aiops secret rm <target>\ninference-aiops secret migrate                         # import legacy plaintext env (INFERENCE_<T>_TOKEN)\ninference-aiops secret rotate-password                 # re-encrypt under a new master password\n```\n\n## Read commands\n\n```bash\ninference-aiops overview [--target <t>]        # Serve deployments + total replicas + queue backpressure\ninference-aiops serve list                     # Ray Serve deployments + replica counts\ninference-aiops serve status <application> <deployment>   # one deployment's status + replica count\ninference-aiops metrics requests               # TTFT / TPOT / e2e latency + token totals (from vLLM /metrics)\ninference-aiops metrics queue                  # running vs waiting requests (backpressure)\ninference-aiops metrics diagnose               # flagship RCA: ranked cause of a latency spike + the knob to turn\n```\n\n## Write commands (governed; risk tier in parentheses)\n\n```bash\ninference-aiops serve scale <application> <deployment> <num_replicas>   # (med) reversible\ninference-aiops serve scale-to-zero <application> <deployment> [--dry-run]   # (high) --dry-run + double confirm; strands ingress\n```\n\n> The remaining writes — `scale_replicas_down`, `drain_replica`,\n> `autoscale_config_update`, `lora_load` / `lora_unload`, `model_sleep` /\n> `model_wake` (dev-mode servers only),\n> `ray_job_cancel`, `replica_restart`, `model_deploy` / `model_undeploy`,\n> `deployment_redeploy`, `routing_policy_update` — are exposed via the MCP\n> server. High-risk ones support a dry-run preview.\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the default/first target)\n- `--dry-run` — print the API call that would be made, change nothing\n- State-changing commands (e.g. `serve scale-to-zero`) require two confirmations\n\nFile v0.10.4:references/setup-guide.md\n\n# inference-aiops setup & security guide\n\n> Not yet validated against a live cluster — see `docs/VERIFICATION.md`.\n\n## 1. Install\n\n```bash\nuv tool install inference-aiops\n```\n\n## 2. (Optional) create a bearer token\n\nA bearer token is **optional** — many vLLM / Ray stacks run open. Only create\none if your API requires it (e.g. vLLM started with `--api-key`, or an\nauthenticating proxy in front of the Ray dashboard). inference-aiops sends it as\n`Authorization: Bearer <token>` to both the Ray dashboard and vLLM.\n\n## 3. Onboard\n\n```bash\ninference-aiops init\n```\n\nThe wizard collects (non-secret) connection details into\n`~/.inference-aiops/config.yaml`: the **serving engine** (`vllm` / `sglang` /\n`tgi`), a **host**, the engine's **port**, and the **scheme** (http/https). For\nthe vLLM engine it also collects the **Ray dashboard port** (default 8265) — its\ncontrol plane. A bearer token is stored **encrypted** into\n`~/.inference-aiops/secrets.enc` **only if the API requires one**. Example config:\n\n```yaml\ntargets:\n  - name: prod                 # vLLM + Ray Serve control plane\n    host: 10.0.0.20\n    engine: vllm\n    scheme: http\n    ray_port: 8265\n    vllm_port: 8000            # vLLM engine port (legacy key; == engine_port)\n    verify_ssl: false          # self-signed lab certs only\n  - name: sg                    # SGLang (single-process, no Ray)\n    host: 10.0.0.21\n    engine: sglang\n    engine_port: 30000         # SGLang default\n  - name: edge                  # TGI (single-process, no Ray)\n    host: 10.0.0.22\n    engine: tgi\n    engine_port: 8080          # TGI default\n```\n\nDefault engine ports: **vLLM 8000, SGLang 30000, TGI 8080**. `engine_port` (or the\nlegacy `vllm_port`) sets the engine's HTTP port. Only the **vLLM** engine uses\n`ray_port`; SGLang and TGI are single-process servers with no Ray dashboard, so\ntheir scale/drain writes raise a teaching error rather than issuing an impossible\ncontrol-plane call.\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nIf a token is stored, export the master password so the encrypted store unlocks\nwithout a prompt (no token stored → nothing to export):\n\n```bash\nexport INFERENCE_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## 5. Laptop self-test (~80% of the tool, free)\n\nMost of the tool self-tests on a laptop with no cloud GPUs:\n\n- **vLLM** — run on a single GPU, or use a CPU-mock, exposing the OpenAI API and\n  Prometheus `/metrics` (default port 8000).\n- **Ray** — one local head node: `ray start --head` (Ray dashboard on 8265),\n  serving a small Serve app.\n\nPoint a target at `host: 127.0.0.1` with `ray_port: 8265` / `vllm_port: 8000`\nand run `inference-aiops doctor`. Reads, RCA, and most scaling ops exercise\nend-to-end. Still unverified: multi-GPU tensor/pipeline-parallel deployments,\nreal GPU thermal/throttle telemetry, and multi-node drain (see\n`docs/VERIFICATION.md`).\n\n## Credential security (when a token is used)\n\n- The token is **never** written to disk in plaintext. It lives only in\n  `~/.inference-aiops/secrets.enc`, encrypted with Fernet (AES-128-CBC + HMAC),\n  the key derived from your master password via scrypt. Only a per-store random\n  salt and the ciphertext are on disk (chmod 600); the master password is never\n  stored.\n- A legacy plaintext env var `INFERENCE_<TARGET_NAME_UPPER>_TOKEN` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `inference-aiops secret migrate`.\n- The token is held only in memory during a session and is never logged or\n  echoed; exception text and tracebacks are scrubbed of secret-shaped strings\n  before being written to the audit log.\n\n## Optional audit annotations\n\nThe tool does not require an approver — whether a high-risk write (scale-down,\nscale-to-zero, drain, LoRA unload, hot-swap, replica restart, undeploy, redeploy)\nshould happen is the agent's decision or the connecting environment's permission.\nIf you want the audit row to carry who ran a change and why, set these; they are\nrecorded when present and never required:\n\n```bash\nexport INFERENCE_AUDIT_APPROVED_BY='you'\nexport INFERENCE_AUDIT_RATIONALE='off-peak cost save'\n```\n\n## Governance harness state\n\nState lives under `~/.inference-aiops/` (relocate with `INFERENCE_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with the risk tier (a descriptive\n  label, not a gate) and any approver/rationale annotation\n- `undo.db` — inverse descriptors for reversible writes (scale, autoscale-config,\n  routing, hot-swap, LoRA load)\n- budget / runaway guard — a safety backstop (not authorization): caps cumulative\n  tool calls and wall-time; trips on tight poll/retry loops\n\n## Verify\n\n```bash\ninference-aiops doctor\n```\n\n`doctor` checks the config file, the encrypted store and its permissions (if a\ntoken is configured), and — unless `--skip-auth` — connectivity. For a **vLLM**\ntarget it probes the **Ray dashboard** and **vLLM** independently, so a half-up\ncluster is reported precisely; for a **SGLang / TGI** target it probes the\nengine's health endpoint and running-model inventory (no Ray).\n\nFile v0.10.4:skill-card.md\n\n## Description:\n\nInference AIops helps agents operate GPU inference serving stacks across vLLM, Ray Serve, SGLang, and TGI with latency diagnosis, metrics inspection, scaling, drain, LoRA, model lifecycle, GPU utilization, Ray job, and cost-per-token workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and inference operations engineers use this skill to inspect, diagnose, and operate GPU inference clusters, especially vLLM/Ray Serve deployments and SGLang/TGI servers. It supports both read-only observability tasks and state-changing operations such as scaling, draining, LoRA changes, model lifecycle actions, and Ray job cancellation.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can invoke disruptive cluster write actions through MCP without an enforced read-only mode or approval gate.\n\nMitigation: Limit network access and credentials to the intended inference endpoints, prefer read-only connectivity in production, and require operator change-control before write-capable sessions.\n\nRisk: High-impact operations such as scale-to-zero, draining, restarts, undeploys, and model changes can interrupt or strand live traffic.\n\nMitigation: Use dry-run previews where available, verify current traffic and queue state first, and keep rollback or undo steps ready before executing changes.\n\nRisk: Stored configuration and optional secrets under the inference-aiops state directory can grant access to inference targets.\n\nMitigation: Protect ~/.inference-aiops, INFERENCE_AIOPS_MASTER_PASSWORD, INFERENCE_AIOPS_CONFIG, and any legacy token environment variables with production credential controls.\n\n## Reference(s):\n\n- [Inference-AIops GitHub repository](https://github.com/AIops-tools/Inference-AIops)\n- [capabilities.md](references/capabilities.md)\n- [cli-reference.md](references/cli-reference.md)\n- [setup-guide.md](references/setup-guide.md)\n- [agent-guardrails.md](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Markdown, Shell commands, Configuration, API calls]\n\n**Output Format:** [Markdown guidance with inline shell commands and tool-call recommendations]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include operational recommendations, dry-run instructions, configuration steps, and risk-aware next actions for supported inference-serving targets.]\n\n## Skill Version(s):\n\n0.10.4 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.3: 7 files, 19596 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (2729b), SKILL.md (17210b), _meta.json (135b)\n\nFile v0.10.3:SKILL.md\n\n---\nname: inference-aiops\nslug: inference-aiops\ndisplayName: \"Inference AIops\"\nsummary: \"Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Inference-AIops\ntags: [aiops, mcp, governance, inference]\ndescription: >\n  Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens.\n  Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster.\n  Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray).\n  Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).\ninstaller:\n  kind: uv\n  package: inference-aiops\nargument-hint: \"[deployment/model name or describe your inference-cluster task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"inference-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"INFERENCE_AIOPS_CONFIG\",\"INFERENCE_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Inference-AIops\",\"emoji\":\"🚀\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.inference-aiops/ (relocatable via INFERENCE_AIOPS_HOME).\n  Auth: a bearer token is OPTIONAL — many vLLM / Ray stacks run open. When the API requires one it is stored ENCRYPTED in ~/.inference-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'inference-aiops init' to onboard, or 'inference-aiops secret set <target>' to add one. The store is unlocked by a master password from INFERENCE_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var INFERENCE_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback (migrate with 'inference-aiops secret migrate'). The token is sent as an Authorization: Bearer header at request time and held only in memory; it is never logged or echoed.\n  State-changing operations require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + risk-tier label — it records, not authorizes). The fragile prod ops — scale_replicas_down, scale_to_zero, drain_replica, lora_unload, model_sleep, replica_restart, model_undeploy, deployment_redeploy — are high-risk with a dry_run preview; reversible writes (scale, autoscale-config, routing, sleep, LoRA load) record an undo descriptor.\n  Engines: vLLM (with its Ray Serve control plane), SGLang, and TGI. SGLang/TGI are single-process servers with engine-agnostic observability (health, running-model inventory, request metrics, queue depth, latency RCA); Ray-shaped scale/drain writes are vLLM-only and raise a teaching error on a SGLang/TGI target.\n  Metrics: each engine's Prometheus /metrics endpoint is parsed directly — no Prometheus server is required.\n  Webhooks: none — no outbound calls beyond the configured Ray dashboard and vLLM services.\n  SSL: verify_ssl defaults to true; disable only for self-signed lab certificates.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked vLLM/Ray responses; unverified against multi-GPU tensor/pipeline-parallel deployments, real GPU thermal/throttle telemetry, and multi-node drain (see docs/VERIFICATION.md).\n---\n\n# Inference AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor.** Product and trademark names belong to their owners. Source at [github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops) under the MIT license.\n\nGoverned GPU-inference operations for **vLLM** (OpenAI API + Prometheus `/metrics`) and **Ray Serve / Ray Jobs** (Ray dashboard), plus the single-process serving engines **SGLang** and **TGI** — **39 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.inference-aiops/`, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels on every audit row. The flagship `diagnose_latency_spike` folds queue depth + KV-cache pressure + prefix-cache locality into a ranked cause and the specific knob to turn; the engine-agnostic `diagnose_engine_latency` does the same across whatever signals SGLang/TGI expose. Each engine's Prometheus `/metrics` is parsed directly — **no Prometheus server required**.\n\n> **Standalone**: the governance harness is bundled in the package (`inference_aiops.governance`) — no external skill-family dependency. A bearer token is **optional** (many stacks run open).\n\n## What This Skill Does\n\n| Group | Tools | Count | Read or Write |\n|-------|-------|:-----:|:-------------:|\n| **Metrics & RCA** (vLLM) | request metrics, queue depth, KV-cache stats, diagnose latency spike, diagnose low utilisation | 5 | 5 read |\n| **Engine-agnostic** (vLLM/SGLang/TGI) | engine health, engine inventory, engine request metrics, engine queue depth, diagnose engine latency | 5 | 5 read |\n| **Ray Serve (read)** | deployment list, deployment status, replica list, autoscale config get | 4 | 4 read |\n| **Ray Serve (write)** | scale up (med), scale down (high), scale-to-zero (high), autoscale config update (med), drain replica (high) | 5 | 5 write |\n| **Models / vLLM** | model list, model info, LoRA load (med), LoRA unload (high), base hot-swap (high) | 5 | 2 read / 3 write |\n| **Ray cluster / jobs / GPU** | cluster resources, dashboard status, job list, GPU utilisation, job cancel (med), replica restart (high) | 6 | 4 read / 2 write |\n| **Deploy lifecycle** | deploy (med), undeploy (high), redeploy (high), routing policy update (med) | 4 | 4 write |\n| **Cost** | cost per token | 1 | 1 read |\n\n**23 read, 16 write**, plus `undo_list` / `undo_apply` — **39 MCP tools** in total. The high-risk writes support `dry_run` + double-confirm; reversible writes record an undo descriptor. The engine-agnostic reads cover any engine; the Ray Serve / cluster / deploy write groups are vLLM-only and teach-and-refuse on a SGLang/TGI target (single-process engines have no Ray control plane).\n\n## Quick Install\n\n```bash\nuv tool install inference-aiops\ninference-aiops init       # interactive wizard: engine (vllm/sglang/tgi) + host + port + scheme (token optional)\ninference-aiops doctor     # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/inference-aiops\nopenclaw skills info inference-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Triage a cluster (`overview`): Serve deployments, total replicas, queue backpressure\n- Diagnose slow inference (`metrics diagnose` / `diagnose_latency_spike`): rank the cause (queue depth vs KV-cache preemption vs prefix-cache locality) and get the knob to turn\n- Find idle GPUs and over-provisioned replicas (`diagnose_low_utilization`)\n- Scale a Ray Serve deployment up/down, **scale-to-zero** to stop cost bleed, or update autoscale bounds\n- **Drain** a replica gracefully before a node reboot (finishes in-flight requests)\n- Load/unload a **LoRA** adapter; **hot-swap** a base model (Sleep-Mode swap, captures the prior model)\n- Inspect GPU utilisation per node, list/cancel Ray jobs, restart a stuck replica\n- Compute **cost per million tokens** from throughput × GPU $/hr\n- Observe an **SGLang** or **TGI** server (`engine_health`, `engine_inventory`, `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) — single-process engines with no Ray control plane\n\n**Do NOT use for** non-inference infrastructure (hypervisors, storage appliances, backup products, general container workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools. This skill is scoped to GPU inference serving (vLLM + Ray).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| vLLM / Ray Serve inference: latency RCA, autoscale, drain, LoRA, cost/token | **inference-aiops** (this skill) |\n| SGLang / TGI serving: health, running-model inventory, request metrics, queue depth, latency RCA | **inference-aiops** (this skill — engine-agnostic reads) |\n| Any non-inference infrastructure (hypervisor, storage, backup, general clusters, network, OT) | the appropriate **other AIops-tools** line |\n\n## Common Workflows\n\n### 1. \"Inference got slow this afternoon\" (flagship RCA → the right knob)\n\n1. `inference-aiops doctor` → confirm the vLLM endpoint and Ray dashboard are actually reachable before blaming the model\n2. `inference-aiops overview` → Serve deployments, total replicas, and whether queue backpressure is cluster-wide or one deployment\n3. `inference-aiops metrics diagnose` (MCP: `diagnose_latency_spike`) → a **ranked** cause with the measured numbers: is `waiting` queue depth high (backpressure)? Are there KV-cache **preemptions** (`kv_cache_stats`)? Has the **prefix-cache hit rate** dropped (routing lost locality)?\n4. Turn the knob the RCA names, not a guess:\n   - backpressure → `inference-aiops serve scale <app> <deployment> --replicas N` (`scale_replicas_up`, reversible, prior count captured)\n   - KV-cache preemption → `autoscale_config_update` to lower the concurrent-request cap (reversible, prior config captured)\n   - lost locality → `routing_policy_update` to prefix-aware / session-affinity (reversible)\n5. Re-check `inference-aiops metrics requests` (TTFT / TPOT / e2e) and `inference-aiops metrics queue` to confirm the p99 actually moved\n6. **Failure branch**: if the fix makes it worse, `inference-aiops undo list` → `inference-aiops undo apply <id>` restores the exact prior replica count / autoscale config / routing policy. If `diagnose_latency_spike` reports no clear cause, the bottleneck is likely upstream of serving — check `gpu_utilization` for a throttling or shared-GPU problem before scaling anything.\n\n### 2. Off-peak cost save: scale a deployment down to zero and bring it back\n\n1. `inference-aiops metrics requests` → confirm traffic really is idle, not just briefly quiet\n2. `diagnose_low_utilization` → the deployments actually burning GPU for nothing, with the measured utilisation\n3. `cost_per_token` → quantify the bleed ($/1M tokens at the current throughput) so the change is justifiable in the audit trail\n4. (optional) `export INFERENCE_AUDIT_APPROVED_BY=you INFERENCE_AUDIT_RATIONALE=\"off-peak cost save\"` → annotates the audit row with who/why; recorded when set, never required\n5. `inference-aiops serve scale-to-zero <app> <deployment> --dry-run`, then re-run without `--dry-run` → **high** risk, double confirmation. `scale_to_zero` stops the bleed but **strands ingress** — requests will queue or fail until replicas return\n6. To restore: `inference-aiops undo apply <id>` (replays the captured prior replica count) or `inference-aiops serve scale <app> <deployment> --replicas N`\n7. **Failure branch**: if traffic arrives while at zero, restore immediately via undo — do not wait for autoscale, since `scale_to_zero` may have been applied outside the autoscaler's floor. If the restore fails, `serve status` will show the deployment unhealthy; `deployment_redeploy` is the last resort (high risk, disruptive).\n\n### 3. Drain a replica before a node reboot\n\n1. `inference-aiops serve list` / `replica_list` → identify the replicas pinned to the node you are about to reboot\n2. `queue_depth` → confirm the remaining replicas can absorb the load; if not, `scale_replicas_up` **first** so draining does not cause a brownout\n3. `drain_replica <app> <deployment> <replica_id> --dry-run`, then confirm → **high** risk; the drain finishes in-flight requests before removing the replica\n4. Watch `replica_list` until the replica is gone and `request_metrics` shows no error spike, then reboot the node\n5. **Failure branch**: if the drain hangs on a long-running request, `replica_restart` forcibly cycles it — that **drops** in-flight requests, so only reach for it once you accept the loss. Multi-node drain has not been verified against a live cluster (see `docs/VERIFICATION.md`).\n\n### 4. Free GPU memory between bursts with Sleep Mode, then resume\n\n1. `model_is_sleeping` → is the engine already suspended? `null` means the engine did not report it — that is UNKNOWN, not awake, so resolve it before writing\n2. `request_metrics` / `queue_depth` → confirm the engine is actually idle; sleeping a busy engine drops live traffic\n3. `model_sleep --dry-run`, then confirm → **high** risk. Level 1 offloads the weights to CPU RAM and wakes fast; level 2 discards them, so waking reloads from disk. The undo descriptor is recorded **only** if the engine was observed awake first — an already-sleeping engine records none, so an undo can never wake something this call did not suspend\n4. Verify: `model_is_sleeping` reports true, and GPU memory has been released (`gpu_utilization`)\n5. Resume with `model_wake` (medium risk), or `inference-aiops undo apply <id>` to replay the recorded inverse. `model_wake` itself records **no** undo: vLLM reports whether the engine sleeps but never at which level, and guessing between level 1 and level 2 would be inventing a prior state\n6. **Failure branch**: if any of the three tools reports that the route does not exist, the server was **not** started with `VLLM_SERVER_DEV_MODE=1`. That is a server start-up flag, not a fault in the tool and not a stale id — restart vLLM with the flag, or leave Sleep Mode off if this is a production deployment that should not expose it.\n\n> vLLM has **no** in-place base-model swap. Sleep Mode suspends and resumes the *same* model; serving a different base model means restarting vLLM with a different `--model`. For adapter-level changes use `lora_load` (reversible) and `lora_unload` (high).\n\n## Governance & Safety\n\nThe skill delivers reads and writes and records them; it does **not** decide\nwhether a write is permitted. That is your agent's judgement, or the permission\nof the environment you connect it with (a network path that only reaches the\nread/metrics endpoints, a Ray dashboard without its job-submission API — writes\nthen fail at the server). There is no read-only switch, policy file, or approval\ngate.\n\n- **Audit is the guarantee, and it is not bypassable.** Every operation — MCP and CLI alike — is logged to `~/.inference-aiops/audit.db` (relocatable via `INFERENCE_AIOPS_HOME`): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.\n- `INFERENCE_AUDIT_APPROVED_BY` / `INFERENCE_AUDIT_RATIONALE` are optional annotations recorded on the audit row (who/why); they are never required and never block.\n- **Runaway guard** — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker.\n- The fragile prod writes support `--dry-run` / `dry_run=True` and double confirmation at the CLI.\n- Reversible writes (scale, autoscale-config, routing, hot-swap, LoRA load) capture before-state and record an inverse descriptor.\n\n## References\n\n- `references/capabilities.md` — full tool → backend → endpoint → returns reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, optional token, and connectivity\n\nFile v0.10.3:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"inference-aiops\",\n  \"version\": \"0.10.3\",\n  \"publishedAt\": 1789452071929\n}\n\nFile v0.10.3:references/agent-guardrails.md\n\n# Agent guardrails — running inference-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the environment's. The tool\ndoes not gate it — there is no read-only switch and no approval prompt to\nconfigure. The two right places to control read vs write:\n\n- **The environment you connect it to.** Restrict the network path so the tool\n  can only reach the read/metrics endpoints, or run the Ray dashboard without its\n  job-submission API. A write then fails at the server, which is the only place\n  the permission actually lives — no skill-side flag can be argued around by a\n  model, but a blocked endpoint cannot.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the engine or Ray dashboard did not return comes back as `null`, never as `\"\"`. An absent job `entrypoint`, a model's `parent` adapter, a replica `state`, or a server-info `version` is distinguishable from an empty one. |\n| \"Tell me if the output was cut off\" | `ray_job_list` returns `{\"jobs\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`. Truncation is measured against the full fetch, not guessed from a length coincidence. |\n| \"Say when a metric isn't available\" | Signals the engine does not expose come back as `null` rather than `0`. SGLang and TGI expose fewer metrics than vLLM; `diagnose_engine_latency` skips a signal it cannot read instead of fabricating it, and `signalsChecked` shows exactly what it looked at. |\n| \"Don't suggest scaling on an engine that can't scale\" | Multi-replica scale / drain / autoscale are Ray Serve control-plane actions. On a single-process engine (SGLang, TGI) those tools raise `EngineCapabilityError` with an explanation, rather than issuing a call that could never succeed. |\n| \"Confirm before anything disruptive\" | Traffic-affecting operations (`model_undeploy`, `deployment_redeploy`, `scale_to_zero`, `drain_replica`, `replica_restart`, `lora_unload`, `model_sleep`) require a `--dry-run`-able preview + double confirmation at the CLI. |\n| \"Log what you did\" | Every call is audited to `~/.inference-aiops/audit.db` regardless of what the model says it did. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a GPU inference cluster through the inference-aiops MCP tools\n(vLLM / SGLang / TGI serving engines, plus a Ray Serve control plane).\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a tool.\n  Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"healthy\".\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null metric means the engine does not expose that signal. Report it as \"not\n  available\" — never substitute 0, and never compare a null against a threshold.\n- Report values exactly as returned. Do not normalise or prettify model ids,\n  deployment names, replica states, or Ray job statuses.\n- When diagnose_engine_latency or diagnose_latency_spike returns probableCauses,\n  work through them in the order given and cite the measured number in each\n  cause's \"signal\" — do not substitute your own theory of the bottleneck.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a latency, throughput, or capacity problem unless a tool result\n  supports it. High GPU utilisation is not by itself a fault.\n- Do not confuse the identifier kinds: a Ray *application* name, a *deployment*\n  name within it, a *replica* id, a Ray *job* id (raysubmit_…), and a served\n  *model* id are four different things. Never pass one where another is expected.\n- cost_per_token is arithmetic over a price you supply, not a billing figure.\n  Present it as an estimate with its inputs.\n```\n\n## Recommended setup for a local model\n\nStart with a path that *cannot* write — restrict the network route to the\nread/metrics endpoints, or expose the Ray dashboard without its job-submission\nAPI — verify, and widen access only when you trust the setup. The\ntraffic-affecting operations here (`scale_to_zero`, `drain_replica`,\n`model_undeploy`) strand or drop live requests and are cheap to invoke:\n\n```bash\ninference-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport INFERENCE_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport INFERENCE_AUDIT_RATIONALE=\"scaling llm-app down for the maintenance window\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the `diagnose_*` tools —\n  `diagnose_engine_latency`, `diagnose_latency_spike`, `diagnose_low_utilization`\n  do the multi-signal correlation inside one call, so the model does not have to\n  chain reads and keep deployment/replica ids straight.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions and use `limit` deliberately rather than dumping a cluster's whole\n  job history.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.3:references/capabilities.md\n\n# inference-aiops capabilities\n\n> 39 MCP tools (23 read, 16 write, 2 undo). Serving engines:\n> **vLLM** (OpenAI API + Prometheus `/metrics`, default 8000) with its **Ray**\n> dashboard control plane (Serve + Jobs, default 8265), plus the single-process\n> engines **SGLang** (OpenAI API + `/get_server_info` + Prometheus `/metrics`,\n> default 30000) and **TGI** (`/info` + Prometheus `/metrics`, default 8080).\n> Endpoints modelled against those APIs; need live verification.\n\n## Metrics & RCA — vLLM (read, 5)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `request_metrics` | vLLM | `GET /metrics` | TTFT, TPOT, e2e latency (avg/p50/p90/p99), prompt/generation token totals, request counts |\n| `queue_depth` | vLLM | `GET /metrics` | running vs waiting requests (backpressure), scheduler state |\n| `kv_cache_stats` | vLLM | `GET /metrics` | KV-cache utilisation %, prefix-cache hit rate, preemption count |\n| `diagnose_latency_spike` | vLLM | `GET /metrics` (fold) | **ranked cause** (queue backpressure / KV-cache preemption / prefix-cache locality) + the specific knob to turn |\n| `diagnose_low_utilization` | vLLM | `GET /metrics` (fold) | idle-GPU / over-provisioned / routing-stranded diagnosis + what to scale down |\n\n## Engine-agnostic — vLLM / SGLang / TGI (read, 5)\n\nWork against **any** supported engine, reading each engine's own paths and metric\nnames (vLLM `vllm:*`, SGLang `sglang:*`, TGI `tgi_*`). A signal an engine does not\nexpose (e.g. TGI has no TTFT or KV-cache metric) degrades to `null` rather than\nbeing guessed.\n\n| Tool | Endpoint(s) | Returns |\n|------|-------------|---------|\n| `engine_health` | `GET /health` | engine liveness (`healthy` bool) + engine label |\n| `engine_inventory` | `GET /v1/models` (vLLM/SGLang) or `/info` (TGI); `/get_server_info` (SGLang) | running-model id(s) + best-effort server info (model, version, max concurrency) |\n| `engine_request_metrics` | `GET /metrics` | TTFT / TPOT / e2e latency + generation-token totals, per engine's exposition (null where unexposed) |\n| `engine_queue_depth` | `GET /metrics` | running vs waiting requests + backpressure flag (SGLang `num_queue_reqs`, TGI `tgi_queue_size`) |\n| `diagnose_engine_latency` | `GET /metrics` (fold) | **ranked cause** across the signals the engine exposes (queue backpressure / KV-token-cache pressure / cache locality) + the knob to turn |\n\n## Ray Serve — read (4)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `serve_deployment_list` | Ray | `GET /api/serve/applications/` | all Serve deployments: status, replica count, target |\n| `deployment_status` | Ray | `GET /api/serve/applications/` | one deployment's status + current/target replica count |\n| `replica_list` | Ray | `GET /api/serve/applications/` | per-replica id, state, node |\n| `autoscale_config_get` | Ray | `GET /api/serve/applications/` | min/max replicas, target ongoing requests |\n\n## Ray Serve — write (5)\n\n| Tool | Risk | Backend | Endpoint | Undo / safety |\n|------|------|---------|----------|---------------|\n| `scale_replicas_up` | med | Ray | `PUT /api/serve/applications/{app}/deployments/{dep}` | reversible (records prior count) |\n| `scale_replicas_down` | **high** | Ray | `PUT …/deployments/{dep}` | dry-run; captures prior count → undo |\n| `scale_to_zero` | **high** | Ray | `PUT …/deployments/{dep}` | dry-run; stops cost bleed but **strands ingress**; captures prior count → undo |\n| `autoscale_config_update` | med | Ray | `PUT …/deployments/{dep}/autoscale` | reversible (records prior bounds) |\n| `drain_replica` | **high** | Ray | `POST …/deployments/{dep}` (drain) | dry-run; graceful — finishes in-flight requests; no undo |\n\n## Models / vLLM (5)\n\n| Tool | R/W (risk) | Backend | Endpoint | Notes |\n|------|-----------|---------|----------|-------|\n| `model_list` | read | vLLM | `GET /v1/models` | served model ids |\n| `model_info` | read | vLLM | `GET /v1/models` | one model's detail (normalised) |\n| `model_is_sleeping` | read | vLLM | `GET /is_sleeping` | dev-mode only; `null` = engine did not say (UNKNOWN, not awake) |\n| `lora_load` | write (med) | vLLM | `POST /v1/load_lora_adapter` | reversible (undo unloads it) |\n| `lora_unload` | write (**high**) | vLLM | `POST /v1/unload_lora_adapter` | dry-run |\n| `model_sleep` | write (**high**) | vLLM | `POST /sleep?level=N` | dev-mode only; dry-run; level 1 offloads weights to CPU RAM, 2 discards them; captures `wasSleeping` → undo wakes only what it suspended |\n| `model_wake` | write (med) | vLLM | `POST /wake_up` | dev-mode only; dry-run; records **no** undo — vLLM never reports the prior sleep *level*, so re-sleeping would be a guess |\n\n> **Sleep Mode needs `VLLM_SERVER_DEV_MODE=1`.** vLLM registers `/sleep`, `/wake_up`\n> and `/is_sleeping` only under that flag, so on a normal production server these\n> three tools report that the route does not exist and why — a 404 here means the\n> server was not started in dev mode, not that an id was stale. Sleep Mode suspends\n> the **same** model; vLLM has no in-place base-model swap (that needs a restart\n> with a different `--model`).\n\n## Ray cluster / jobs / GPU (6)\n\n| Tool | R/W (risk) | Backend | Endpoint | Returns / notes |\n|------|-----------|---------|----------|-----------------|\n| `ray_cluster_resources` | read | Ray | `GET /api/cluster_status` | CPU/GPU total vs allocated |\n| `ray_dashboard_status` | read | Ray | Ray dashboard | dashboard reachability/version |\n| `ray_job_list` | read | Ray | `GET /api/jobs/` | submitted jobs + status |\n| `gpu_utilization` | read | Ray | `GET /api/nodes` | per-node GPU count, utilisation %, memory (best-effort) |\n| `ray_job_cancel` | write (med) | Ray | `POST /api/jobs/{id}/stop` | cancel a running job |\n| `replica_restart` | write (**high**) | Ray | `GET/PUT /api/serve/applications/` | dry-run; restart a stuck replica |\n\n## Deploy lifecycle (4, write)\n\n| Tool | Risk | Backend | Endpoint | Undo / safety |\n|------|------|---------|----------|---------------|\n| `model_deploy` | med | Ray | `PUT /api/serve/applications/` | deploy an application |\n| `model_undeploy` | **high** | Ray | `DELETE/PUT /api/serve/applications/` | dry-run |\n| `deployment_redeploy` | **high** | Ray | `PUT /api/serve/applications/` | dry-run |\n| `routing_policy_update` | med | Ray | `PUT /api/serve/applications/` | reversible; prefix-aware / session-affinity routing to fix cache locality |\n\n## Cost (read, 1)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `cost_per_token` | vLLM | `GET /metrics` + GPU $/hr | deterministic $/1M tokens from measured throughput × GPU hourly rate |\n\n## SGLang / TGI writes (control-plane teaching error)\n\nSGLang and TGI are **single-process servers** with no Ray Serve control plane, so\nthe Ray-shaped write groups above (scale / drain / autoscale / deploy / redeploy /\nrouting / job-cancel / replica-restart) do not apply to them. Attempting one\nagainst a SGLang/TGI target raises `EngineCapabilityError` with a teaching message\npointing at a real horizontal-scale layer (Ray Serve / Kubernetes / a load\nbalancer). Their supported surface is the **engine-agnostic read** group above.\n\n## Out of scope (by design)\n\n- Cluster **provisioning** (spinning up GPU nodes, driver/CUDA install)\n- vLLM/Ray **install or version upgrades**\n- Multi-node **drain / reboot orchestration** (single-replica drain only, unverified at multi-node scale)\n- Non-inference infrastructure (use the appropriate other AIops-tools line)\n\nWant one of these? Open an issue or PR — feedback and contributions welcome.\n\nFile v0.10.3:references/cli-reference.md\n\n# inference-aiops CLI reference\n\n> Serving engines: vLLM (OpenAI API + Prometheus `/metrics`,\n> default 8000) with its Ray dashboard control plane (Serve + Jobs, default 8265),\n> plus single-process SGLang (default 30000) and TGI (default 8080); endpoints\n> need live verification.\n>\n> The CLI is a convenience subset. The full 35-tool surface — including the\n> engine-agnostic reads (`engine_health`, `engine_inventory`,\n> `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) that\n> cover SGLang/TGI — is via the MCP server (`inference-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\ninference-aiops init                      # interactive wizard: engine (vllm/sglang/tgi) + host + port\ninference-aiops doctor [--skip-auth]      # config + secret store + connectivity — vLLM: Ray + vLLM; SGLang/TGI: engine health + inventory\ninference-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.inference-aiops/secrets.enc — only if a token is used)\n\n```bash\ninference-aiops secret set <target> [--value <token>]  # store a bearer token (hidden prompt if no --value)\ninference-aiops secret list                            # names only — values never shown\ninference-aiops secret rm <target>\ninference-aiops secret migrate                         # import legacy plaintext env (INFERENCE_<T>_TOKEN)\ninference-aiops secret rotate-password                 # re-encrypt under a new master password\n```\n\n## Read commands\n\n```bash\ninference-aiops overview [--target <t>]        # Serve deployments + total replicas + queue backpressure\ninference-aiops serve list                     # Ray Serve deployments + replica counts\ninference-aiops serve status <application> <deployment>   # one deployment's status + replica count\ninference-aiops metrics requests               # TTFT / TPOT / e2e latency + token totals (from vLLM /metrics)\ninference-aiops metrics queue                  # running vs waiting requests (backpressure)\ninference-aiops metrics diagnose               # flagship RCA: ranked cause of a latency spike + the knob to turn\n```\n\n## Write commands (governed; risk tier in parentheses)\n\n```bash\ninference-aiops serve scale <application> <deployment> <num_replicas>   # (med) reversible\ninference-aiops serve scale-to-zero <application> <deployment> [--dry-run]   # (high) --dry-run + double confirm; strands ingress\n```\n\n> The remaining writes — `scale_replicas_down`, `drain_replica`,\n> `autoscale_config_update`, `lora_load` / `lora_unload`, `model_sleep` /\n> `model_wake` (dev-mode servers only),\n> `ray_job_cancel`, `replica_restart`, `model_deploy` / `model_undeploy`,\n> `deployment_redeploy`, `routing_policy_update` — are exposed via the MCP\n> server. High-risk ones support a dry-run preview.\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the default/first target)\n- `--dry-run` — print the API call that would be made, change nothing\n- State-changing commands (e.g. `serve scale-to-zero`) require two confirmations\n\nFile v0.10.3:references/setup-guide.md\n\n# inference-aiops setup & security guide\n\n> Not yet validated against a live cluster — see `docs/VERIFICATION.md`.\n\n## 1. Install\n\n```bash\nuv tool install inference-aiops\n```\n\n## 2. (Optional) create a bearer token\n\nA bearer token is **optional** — many vLLM / Ray stacks run open. Only create\none if your API requires it (e.g. vLLM started with `--api-key`, or an\nauthenticating proxy in front of the Ray dashboard). inference-aiops sends it as\n`Authorization: Bearer <token>` to both the Ray dashboard and vLLM.\n\n## 3. Onboard\n\n```bash\ninference-aiops init\n```\n\nThe wizard collects (non-secret) connection details into\n`~/.inference-aiops/config.yaml`: the **serving engine** (`vllm` / `sglang` /\n`tgi`), a **host**, the engine's **port**, and the **scheme** (http/https). For\nthe vLLM engine it also collects the **Ray dashboard port** (default 8265) — its\ncontrol plane. A bearer token is stored **encrypted** into\n`~/.inference-aiops/secrets.enc` **only if the API requires one**. Example config:\n\n```yaml\ntargets:\n  - name: prod                 # vLLM + Ray Serve control plane\n    host: 10.0.0.20\n    engine: vllm\n    scheme: http\n    ray_port: 8265\n    vllm_port: 8000            # vLLM engine port (legacy key; == engine_port)\n    verify_ssl: false          # self-signed lab certs only\n  - name: sg                    # SGLang (single-process, no Ray)\n    host: 10.0.0.21\n    engine: sglang\n    engine_port: 30000         # SGLang default\n  - name: edge                  # TGI (single-process, no Ray)\n    host: 10.0.0.22\n    engine: tgi\n    engine_port: 8080          # TGI default\n```\n\nDefault engine ports: **vLLM 8000, SGLang 30000, TGI 8080**. `engine_port` (or the\nlegacy `vllm_port`) sets the engine's HTTP port. Only the **vLLM** engine uses\n`ray_port`; SGLang and TGI are single-process servers with no Ray dashboard, so\ntheir scale/drain writes raise a teaching error rather than issuing an impossible\ncontrol-plane call.\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nIf a token is stored, export the master password so the encrypted store unlocks\nwithout a prompt (no token stored → nothing to export):\n\n```bash\nexport INFERENCE_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## 5. Laptop self-test (~80% of the tool, free)\n\nMost of the tool self-tests on a laptop with no cloud GPUs:\n\n- **vLLM** — run on a single GPU, or use a CPU-mock, exposing the OpenAI API and\n  Prometheus `/metrics` (default port 8000).\n- **Ray** — one local head node: `ray start --head` (Ray dashboard on 8265),\n  serving a small Serve app.\n\nPoint a target at `host: 127.0.0.1` with `ray_port: 8265` / `vllm_port: 8000`\nand run `inference-aiops doctor`. Reads, RCA, and most scaling ops exercise\nend-to-end. Still unverified: multi-GPU tensor/pipeline-parallel deployments,\nreal GPU thermal/throttle telemetry, and multi-node drain (see\n`docs/VERIFICATION.md`).\n\n## Credential security (when a token is used)\n\n- The token is **never** written to disk in plaintext. It lives only in\n  `~/.inference-aiops/secrets.enc`, encrypted with Fernet (AES-128-CBC + HMAC),\n  the key derived from your master password via scrypt. Only a per-store random\n  salt and the ciphertext are on disk (chmod 600); the master password is never\n  stored.\n- A legacy plaintext env var `INFERENCE_<TARGET_NAME_UPPER>_TOKEN` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `inference-aiops secret migrate`.\n- The token is held only in memory during a session and is never logged or\n  echoed; exception text and tracebacks are scrubbed of secret-shaped strings\n  before being written to the audit log.\n\n## Optional audit annotations\n\nThe tool does not require an approver — whether a high-risk write (scale-down,\nscale-to-zero, drain, LoRA unload, hot-swap, replica restart, undeploy, redeploy)\nshould happen is the agent's decision or the connecting environment's permission.\nIf you want the audit row to carry who ran a change and why, set these; they are\nrecorded when present and never required:\n\n```bash\nexport INFERENCE_AUDIT_APPROVED_BY='you'\nexport INFERENCE_AUDIT_RATIONALE='off-peak cost save'\n```\n\n## Governance harness state\n\nState lives under `~/.inference-aiops/` (relocate with `INFERENCE_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with the risk tier (a descriptive\n  label, not a gate) and any approver/rationale annotation\n- `undo.db` — inverse descriptors for reversible writes (scale, autoscale-config,\n  routing, hot-swap, LoRA load)\n- budget / runaway guard — a safety backstop (not authorization): caps cumulative\n  tool calls and wall-time; trips on tight poll/retry loops\n\n## Verify\n\n```bash\ninference-aiops doctor\n```\n\n`doctor` checks the config file, the encrypted store and its permissions (if a\ntoken is configured), and — unless `--skip-auth` — connectivity. For a **vLLM**\ntarget it probes the **Ray dashboard** and **vLLM** independently, so a half-up\ncluster is reported precisely; for a **SGLang / TGI** target it probes the\nengine's health endpoint and running-model inventory (no Ray).\n\nFile v0.10.3:skill-card.md\n\n## Description:\n\nInference AIops helps agents observe and operate GPU inference services across vLLM, Ray Serve, SGLang, and TGI, including metrics, latency root-cause analysis, scaling, draining, model operations, and cost estimates.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and inference platform engineers use this skill to inspect GPU inference clusters, diagnose latency and utilization issues, and propose or run operational changes such as scaling, draining, model lifecycle actions, and cost analysis.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill exposes disruptive production write actions without an enforceable local read-only or approval gate.\n\nMitigation: Use a dedicated low-privilege account, keep network access scoped to read-only endpoints unless writes are explicitly needed, require server-side authorization, and require review before high-risk writes.\n\nRisk: Operational credentials may be exposed or misused if tokens are sent over insecure transport or overly broad access paths.\n\nMitigation: Avoid HTTP or disabled TLS verification when using tokens, use separate credentials for Ray control-plane and inference endpoints where possible, and keep secrets in the configured encrypted store.\n\nRisk: Unpinned or unverified package installation can introduce supply-chain uncertainty.\n\nMitigation: Pin and verify the package version before deployment and install only when the environment intentionally grants the agent inference-operations authority.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/inference-aiops)\n- [Project homepage](https://github.com/AIops-tools/Inference-AIops)\n- [Capabilities reference](references/capabilities.md)\n- [Setup and security guide](references/setup-guide.md)\n- [Agent guardrails](references/agent-guardrails.md)\n- [CLI reference](references/cli-reference.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown or plain text with command snippets and structured tool results]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include operational recommendations, dry-run guidance, and configuration steps based on live tool results.]\n\n## Skill Version(s):\n\n0.10.3 (source: evidence.release.version)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.2: 7 files, 19570 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (2755b), SKILL.md (17210b), _meta.json (135b)\n\nFile v0.10.2:SKILL.md\n\n---\nname: inference-aiops\nslug: inference-aiops\ndisplayName: \"Inference AIops\"\nsummary: \"Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Inference-AIops\ntags: [aiops, mcp, governance, inference]\ndescription: >\n  Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens.\n  Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster.\n  Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray).\n  Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).\ninstaller:\n  kind: uv\n  package: inference-aiops\nargument-hint: \"[deployment/model name or describe your inference-cluster task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"inference-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"INFERENCE_AIOPS_CONFIG\",\"INFERENCE_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Inference-AIops\",\"emoji\":\"🚀\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.inference-aiops/ (relocatable via INFERENCE_AIOPS_HOME).\n  Auth: a bearer token is OPTIONAL — many vLLM / Ray stacks run open. When the API requires one it is stored ENCRYPTED in ~/.inference-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'inference-aiops init' to onboard, or 'inference-aiops secret set <target>' to add one. The store is unlocked by a master password from INFERENCE_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var INFERENCE_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback (migrate with 'inference-aiops secret migrate'). The token is sent as an Authorization: Bearer header at request time and held only in memory; it is never logged or echoed.\n  State-changing operations require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + risk-tier label — it records, not authorizes). The fragile prod ops — scale_replicas_down, scale_to_zero, drain_replica, lora_unload, model_sleep, replica_restart, model_undeploy, deployment_redeploy — are high-risk with a dry_run preview; reversible writes (scale, autoscale-config, routing, sleep, LoRA load) record an undo descriptor.\n  Engines: vLLM (with its Ray Serve control plane), SGLang, and TGI. SGLang/TGI are single-process servers with engine-agnostic observability (health, running-model inventory, request metrics, queue depth, latency RCA); Ray-shaped scale/drain writes are vLLM-only and raise a teaching error on a SGLang/TGI target.\n  Metrics: each engine's Prometheus /metrics endpoint is parsed directly — no Prometheus server is required.\n  Webhooks: none — no outbound calls beyond the configured Ray dashboard and vLLM services.\n  SSL: verify_ssl defaults to true; disable only for self-signed lab certificates.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked vLLM/Ray responses; unverified against multi-GPU tensor/pipeline-parallel deployments, real GPU thermal/throttle telemetry, and multi-node drain (see docs/VERIFICATION.md).\n---\n\n# Inference AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor.** Product and trademark names belong to their owners. Source at [github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops) under the MIT license.\n\nGoverned GPU-inference operations for **vLLM** (OpenAI API + Prometheus `/metrics`) and **Ray Serve / Ray Jobs** (Ray dashboard), plus the single-process serving engines **SGLang** and **TGI** — **39 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.inference-aiops/`, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels on every audit row. The flagship `diagnose_latency_spike` folds queue depth + KV-cache pressure + prefix-cache locality into a ranked cause and the specific knob to turn; the engine-agnostic `diagnose_engine_latency` does the same across whatever signals SGLang/TGI expose. Each engine's Prometheus `/metrics` is parsed directly — **no Prometheus server required**.\n\n> **Standalone**: the governance harness is bundled in the package (`inference_aiops.governance`) — no external skill-family dependency. A bearer token is **optional** (many stacks run open).\n\n## What This Skill Does\n\n| Group | Tools | Count | Read or Write |\n|-------|-------|:-----:|:-------------:|\n| **Metrics & RCA** (vLLM) | request metrics, queue depth, KV-cache stats, diagnose latency spike, diagnose low utilisation | 5 | 5 read |\n| **Engine-agnostic** (vLLM/SGLang/TGI) | engine health, engine inventory, engine request metrics, engine queue depth, diagnose engine latency | 5 | 5 read |\n| **Ray Serve (read)** | deployment list, deployment status, replica list, autoscale config get | 4 | 4 read |\n| **Ray Serve (write)** | scale up (med), scale down (high), scale-to-zero (high), autoscale config update (med), drain replica (high) | 5 | 5 write |\n| **Models / vLLM** | model list, model info, LoRA load (med), LoRA unload (high), base hot-swap (high) | 5 | 2 read / 3 write |\n| **Ray cluster / jobs / GPU** | cluster resources, dashboard status, job list, GPU utilisation, job cancel (med), replica restart (high) | 6 | 4 read / 2 write |\n| **Deploy lifecycle** | deploy (med), undeploy (high), redeploy (high), routing policy update (med) | 4 | 4 write |\n| **Cost** | cost per token | 1 | 1 read |\n\n**23 read, 16 write**, plus `undo_list` / `undo_apply` — **39 MCP tools** in total. The high-risk writes support `dry_run` + double-confirm; reversible writes record an undo descriptor. The engine-agnostic reads cover any engine; the Ray Serve / cluster / deploy write groups are vLLM-only and teach-and-refuse on a SGLang/TGI target (single-process engines have no Ray control plane).\n\n## Quick Install\n\n```bash\nuv tool install inference-aiops\ninference-aiops init       # interactive wizard: engine (vllm/sglang/tgi) + host + port + scheme (token optional)\ninference-aiops doctor     # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/inference-aiops\nopenclaw skills info inference-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Triage a cluster (`overview`): Serve deployments, total replicas, queue backpressure\n- Diagnose slow inference (`metrics diagnose` / `diagnose_latency_spike`): rank the cause (queue depth vs KV-cache preemption vs prefix-cache locality) and get the knob to turn\n- Find idle GPUs and over-provisioned replicas (`diagnose_low_utilization`)\n- Scale a Ray Serve deployment up/down, **scale-to-zero** to stop cost bleed, or update autoscale bounds\n- **Drain** a replica gracefully before a node reboot (finishes in-flight requests)\n- Load/unload a **LoRA** adapter; **hot-swap** a base model (Sleep-Mode swap, captures the prior model)\n- Inspect GPU utilisation per node, list/cancel Ray jobs, restart a stuck replica\n- Compute **cost per million tokens** from throughput × GPU $/hr\n- Observe an **SGLang** or **TGI** server (`engine_health`, `engine_inventory`, `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) — single-process engines with no Ray control plane\n\n**Do NOT use for** non-inference infrastructure (hypervisors, storage appliances, backup products, general container workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools. This skill is scoped to GPU inference serving (vLLM + Ray).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| vLLM / Ray Serve inference: latency RCA, autoscale, drain, LoRA, cost/token | **inference-aiops** (this skill) |\n| SGLang / TGI serving: health, running-model inventory, request metrics, queue depth, latency RCA | **inference-aiops** (this skill — engine-agnostic reads) |\n| Any non-inference infrastructure (hypervisor, storage, backup, general clusters, network, OT) | the appropriate **other AIops-tools** line |\n\n## Common Workflows\n\n### 1. \"Inference got slow this afternoon\" (flagship RCA → the right knob)\n\n1. `inference-aiops doctor` → confirm the vLLM endpoint and Ray dashboard are actually reachable before blaming the model\n2. `inference-aiops overview` → Serve deployments, total replicas, and whether queue backpressure is cluster-wide or one deployment\n3. `inference-aiops metrics diagnose` (MCP: `diagnose_latency_spike`) → a **ranked** cause with the measured numbers: is `waiting` queue depth high (backpressure)? Are there KV-cache **preemptions** (`kv_cache_stats`)? Has the **prefix-cache hit rate** dropped (routing lost locality)?\n4. Turn the knob the RCA names, not a guess:\n   - backpressure → `inference-aiops serve scale <app> <deployment> --replicas N` (`scale_replicas_up`, reversible, prior count captured)\n   - KV-cache preemption → `autoscale_config_update` to lower the concurrent-request cap (reversible, prior config captured)\n   - lost locality → `routing_policy_update` to prefix-aware / session-affinity (reversible)\n5. Re-check `inference-aiops metrics requests` (TTFT / TPOT / e2e) and `inference-aiops metrics queue` to confirm the p99 actually moved\n6. **Failure branch**: if the fix makes it worse, `inference-aiops undo list` → `inference-aiops undo apply <id>` restores the exact prior replica count / autoscale config / routing policy. If `diagnose_latency_spike` reports no clear cause, the bottleneck is likely upstream of serving — check `gpu_utilization` for a throttling or shared-GPU problem before scaling anything.\n\n### 2. Off-peak cost save: scale a deployment down to zero and bring it back\n\n1. `inference-aiops metrics requests` → confirm traffic really is idle, not just briefly quiet\n2. `diagnose_low_utilization` → the deployments actually burning GPU for nothing, with the measured utilisation\n3. `cost_per_token` → quantify the bleed ($/1M tokens at the current throughput) so the change is justifiable in the audit trail\n4. (optional) `export INFERENCE_AUDIT_APPROVED_BY=you INFERENCE_AUDIT_RATIONALE=\"off-peak cost save\"` → annotates the audit row with who/why; recorded when set, never required\n5. `inference-aiops serve scale-to-zero <app> <deployment> --dry-run`, then re-run without `--dry-run` → **high** risk, double confirmation. `scale_to_zero` stops the bleed but **strands ingress** — requests will queue or fail until replicas return\n6. To restore: `inference-aiops undo apply <id>` (replays the captured prior replica count) or `inference-aiops serve scale <app> <deployment> --replicas N`\n7. **Failure branch**: if traffic arrives while at zero, restore immediately via undo — do not wait for autoscale, since `scale_to_zero` may have been applied outside the autoscaler's floor. If the restore fails, `serve status` will show the deployment unhealthy; `deployment_redeploy` is the last resort (high risk, disruptive).\n\n### 3. Drain a replica before a node reboot\n\n1. `inference-aiops serve list` / `replica_list` → identify the replicas pinned to the node you are about to reboot\n2. `queue_depth` → confirm the remaining replicas can absorb the load; if not, `scale_replicas_up` **first** so draining does not cause a brownout\n3. `drain_replica <app> <deployment> <replica_id> --dry-run`, then confirm → **high** risk; the drain finishes in-flight requests before removing the replica\n4. Watch `replica_list` until the replica is gone and `request_metrics` shows no error spike, then reboot the node\n5. **Failure branch**: if the drain hangs on a long-running request, `replica_restart` forcibly cycles it — that **drops** in-flight requests, so only reach for it once you accept the loss. Multi-node drain has not been verified against a live cluster (see `docs/VERIFICATION.md`).\n\n### 4. Free GPU memory between bursts with Sleep Mode, then resume\n\n1. `model_is_sleeping` → is the engine already suspended? `null` means the engine did not report it — that is UNKNOWN, not awake, so resolve it before writing\n2. `request_metrics` / `queue_depth` → confirm the engine is actually idle; sleeping a busy engine drops live traffic\n3. `model_sleep --dry-run`, then confirm → **high** risk. Level 1 offloads the weights to CPU RAM and wakes fast; level 2 discards them, so waking reloads from disk. The undo descriptor is recorded **only** if the engine was observed awake first — an already-sleeping engine records none, so an undo can never wake something this call did not suspend\n4. Verify: `model_is_sleeping` reports true, and GPU memory has been released (`gpu_utilization`)\n5. Resume with `model_wake` (medium risk), or `inference-aiops undo apply <id>` to replay the recorded inverse. `model_wake` itself records **no** undo: vLLM reports whether the engine sleeps but never at which level, and guessing between level 1 and level 2 would be inventing a prior state\n6. **Failure branch**: if any of the three tools reports that the route does not exist, the server was **not** started with `VLLM_SERVER_DEV_MODE=1`. That is a server start-up flag, not a fault in the tool and not a stale id — restart vLLM with the flag, or leave Sleep Mode off if this is a production deployment that should not expose it.\n\n> vLLM has **no** in-place base-model swap. Sleep Mode suspends and resumes the *same* model; serving a different base model means restarting vLLM with a different `--model`. For adapter-level changes use `lora_load` (reversible) and `lora_unload` (high).\n\n## Governance & Safety\n\nThe skill delivers reads and writes and records them; it does **not** decide\nwhether a write is permitted. That is your agent's judgement, or the permission\nof the environment you connect it with (a network path that only reaches the\nread/metrics endpoints, a Ray dashboard without its job-submission API — writes\nthen fail at the server). There is no read-only switch, policy file, or approval\ngate.\n\n- **Audit is the guarantee, and it is not bypassable.** Every operation — MCP and CLI alike — is logged to `~/.inference-aiops/audit.db` (relocatable via `INFERENCE_AIOPS_HOME`): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.\n- `INFERENCE_AUDIT_APPROVED_BY` / `INFERENCE_AUDIT_RATIONALE` are optional annotations recorded on the audit row (who/why); they are never required and never block.\n- **Runaway guard** — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker.\n- The fragile prod writes support `--dry-run` / `dry_run=True` and double confirmation at the CLI.\n- Reversible writes (scale, autoscale-config, routing, hot-swap, LoRA load) capture before-state and record an inverse descriptor.\n\n## References\n\n- `references/capabilities.md` — full tool → backend → endpoint → returns reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, optional token, and connectivity\n\nFile v0.10.2:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"inference-aiops\",\n  \"version\": \"0.10.2\",\n  \"publishedAt\": 1789223106436\n}\n\nFile v0.10.2:references/agent-guardrails.md\n\n# Agent guardrails — running inference-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the environment's. The tool\ndoes not gate it — there is no read-only switch and no approval prompt to\nconfigure. The two right places to control read vs write:\n\n- **The environment you connect it to.** Restrict the network path so the tool\n  can only reach the read/metrics endpoints, or run the Ray dashboard without its\n  job-submission API. A write then fails at the server, which is the only place\n  the permission actually lives — no skill-side flag can be argued around by a\n  model, but a blocked endpoint cannot.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the engine or Ray dashboard did not return comes back as `null`, never as `\"\"`. An absent job `entrypoint`, a model's `parent` adapter, a replica `state`, or a server-info `version` is distinguishable from an empty one. |\n| \"Tell me if the output was cut off\" | `ray_job_list` returns `{\"jobs\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`. Truncation is measured against the full fetch, not guessed from a length coincidence. |\n| \"Say when a metric isn't available\" | Signals the engine does not expose come back as `null` rather than `0`. SGLang and TGI expose fewer metrics than vLLM; `diagnose_engine_latency` skips a signal it cannot read instead of fabricating it, and `signalsChecked` shows exactly what it looked at. |\n| \"Don't suggest scaling on an engine that can't scale\" | Multi-replica scale / drain / autoscale are Ray Serve control-plane actions. On a single-process engine (SGLang, TGI) those tools raise `EngineCapabilityError` with an explanation, rather than issuing a call that could never succeed. |\n| \"Confirm before anything disruptive\" | Traffic-affecting operations (`model_undeploy`, `deployment_redeploy`, `scale_to_zero`, `drain_replica`, `replica_restart`, `lora_unload`, `model_sleep`) require a `--dry-run`-able preview + double confirmation at the CLI. |\n| \"Log what you did\" | Every call is audited to `~/.inference-aiops/audit.db` regardless of what the model says it did. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a GPU inference cluster through the inference-aiops MCP tools\n(vLLM / SGLang / TGI serving engines, plus a Ray Serve control plane).\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a tool.\n  Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"healthy\".\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null metric means the engine does not expose that signal. Report it as \"not\n  available\" — never substitute 0, and never compare a null against a threshold.\n- Report values exactly as returned. Do not normalise or prettify model ids,\n  deployment names, replica states, or Ray job statuses.\n- When diagnose_engine_latency or diagnose_latency_spike returns probableCauses,\n  work through them in the order given and cite the measured number in each\n  cause's \"signal\" — do not substitute your own theory of the bottleneck.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a latency, throughput, or capacity problem unless a tool result\n  supports it. High GPU utilisation is not by itself a fault.\n- Do not confuse the identifier kinds: a Ray *application* name, a *deployment*\n  name within it, a *replica* id, a Ray *job* id (raysubmit_…), and a served\n  *model* id are four different things. Never pass one where another is expected.\n- cost_per_token is arithmetic over a price you supply, not a billing figure.\n  Present it as an estimate with its inputs.\n```\n\n## Recommended setup for a local model\n\nStart with a path that *cannot* write — restrict the network route to the\nread/metrics endpoints, or expose the Ray dashboard without its job-submission\nAPI — verify, and widen access only when you trust the setup. The\ntraffic-affecting operations here (`scale_to_zero`, `drain_replica`,\n`model_undeploy`) strand or drop live requests and are cheap to invoke:\n\n```bash\ninference-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport INFERENCE_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport INFERENCE_AUDIT_RATIONALE=\"scaling llm-app down for the maintenance window\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the `diagnose_*` tools —\n  `diagnose_engine_latency`, `diagnose_latency_spike`, `diagnose_low_utilization`\n  do the multi-signal correlation inside one call, so the model does not have to\n  chain reads and keep deployment/replica ids straight.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions and use `limit` deliberately rather than dumping a cluster's whole\n  job history.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.2:references/capabilities.md\n\n# inference-aiops capabilities\n\n> 39 MCP tools (23 read, 16 write, 2 undo). Serving engines:\n> **vLLM** (OpenAI API + Prometheus `/metrics`, default 8000) with its **Ray**\n> dashboard control plane (Serve + Jobs, default 8265), plus the single-process\n> engines **SGLang** (OpenAI API + `/get_server_info` + Prometheus `/metrics`,\n> default 30000) and **TGI** (`/info` + Prometheus `/metrics`, default 8080).\n> Endpoints modelled against those APIs; need live verification.\n\n## Metrics & RCA — vLLM (read, 5)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `request_metrics` | vLLM | `GET /metrics` | TTFT, TPOT, e2e latency (avg/p50/p90/p99), prompt/generation token totals, request counts |\n| `queue_depth` | vLLM | `GET /metrics` | running vs waiting requests (backpressure), scheduler state |\n| `kv_cache_stats` | vLLM | `GET /metrics` | KV-cache utilisation %, prefix-cache hit rate, preemption count |\n| `diagnose_latency_spike` | vLLM | `GET /metrics` (fold) | **ranked cause** (queue backpressure / KV-cache preemption / prefix-cache locality) + the specific knob to turn |\n| `diagnose_low_utilization` | vLLM | `GET /metrics` (fold) | idle-GPU / over-provisioned / routing-stranded diagnosis + what to scale down |\n\n## Engine-agnostic — vLLM / SGLang / TGI (read, 5)\n\nWork against **any** supported engine, reading each engine's own paths and metric\nnames (vLLM `vllm:*`, SGLang `sglang:*`, TGI `tgi_*`). A signal an engine does not\nexpose (e.g. TGI has no TTFT or KV-cache metric) degrades to `null` rather than\nbeing guessed.\n\n| Tool | Endpoint(s) | Returns |\n|------|-------------|---------|\n| `engine_health` | `GET /health` | engine liveness (`healthy` bool) + engine label |\n| `engine_inventory` | `GET /v1/models` (vLLM/SGLang) or `/info` (TGI); `/get_server_info` (SGLang) | running-model id(s) + best-effort server info (model, version, max concurrency) |\n| `engine_request_metrics` | `GET /metrics` | TTFT / TPOT / e2e latency + generation-token totals, per engine's exposition (null where unexposed) |\n| `engine_queue_depth` | `GET /metrics` | running vs waiting requests + backpressure flag (SGLang `num_queue_reqs`, TGI `tgi_queue_size`) |\n| `diagnose_engine_latency` | `GET /metrics` (fold) | **ranked cause** across the signals the engine exposes (queue backpressure / KV-token-cache pressure / cache locality) + the knob to turn |\n\n## Ray Serve — read (4)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `serve_deployment_list` | Ray | `GET /api/serve/applications/` | all Serve deployments: status, replica count, target |\n| `deployment_status` | Ray | `GET /api/serve/applications/` | one deployment's status + current/target replica count |\n| `replica_list` | Ray | `GET /api/serve/applications/` | per-replica id, state, node |\n| `autoscale_config_get` | Ray | `GET /api/serve/applications/` | min/max replicas, target ongoing requests |\n\n## Ray Serve — write (5)\n\n| Tool | Risk | Backend | Endpoint | Undo / safety |\n|------|------|---------|----------|---------------|\n| `scale_replicas_up` | med | Ray | `PUT /api/serve/applications/{app}/deployments/{dep}` | reversible (records prior count) |\n| `scale_replicas_down` | **high** | Ray | `PUT …/deployments/{dep}` | dry-run; captures prior count → undo |\n| `scale_to_zero` | **high** | Ray | `PUT …/deployments/{dep}` | dry-run; stops cost bleed but **strands ingress**; captures prior count → undo |\n| `autoscale_config_update` | med | Ray | `PUT …/deployments/{dep}/autoscale` | reversible (records prior bounds) |\n| `drain_replica` | **high** | Ray | `POST …/deployments/{dep}` (drain) | dry-run; graceful — finishes in-flight requests; no undo |\n\n## Models / vLLM (5)\n\n| Tool | R/W (risk) | Backend | Endpoint | Notes |\n|------|-----------|---------|----------|-------|\n| `model_list` | read | vLLM | `GET /v1/models` | served model ids |\n| `model_info` | read | vLLM | `GET /v1/models` | one model's detail (normalised) |\n| `model_is_sleeping` | read | vLLM | `GET /is_sleeping` | dev-mode only; `null` = engine did not say (UNKNOWN, not awake) |\n| `lora_load` | write (med) | vLLM | `POST /v1/load_lora_adapter` | reversible (undo unloads it) |\n| `lora_unload` | write (**high**) | vLLM | `POST /v1/unload_lora_adapter` | dry-run |\n| `model_sleep` | write (**high**) | vLLM | `POST /sleep?level=N` | dev-mode only; dry-run; level 1 offloads weights to CPU RAM, 2 discards them; captures `wasSleeping` → undo wakes only what it suspended |\n| `model_wake` | write (med) | vLLM | `POST /wake_up` | dev-mode only; dry-run; records **no** undo — vLLM never reports the prior sleep *level*, so re-sleeping would be a guess |\n\n> **Sleep Mode needs `VLLM_SERVER_DEV_MODE=1`.** vLLM registers `/sleep`, `/wake_up`\n> and `/is_sleeping` only under that flag, so on a normal production server these\n> three tools report that the route does not exist and why — a 404 here means the\n> server was not started in dev mode, not that an id was stale. Sleep Mode suspends\n> the **same** model; vLLM has no in-place base-model swap (that needs a restart\n> with a different `--model`).\n\n## Ray cluster / jobs / GPU (6)\n\n| Tool | R/W (risk) | Backend | Endpoint | Returns / notes |\n|------|-----------|---------|----------|-----------------|\n| `ray_cluster_resources` | read | Ray | `GET /api/cluster_status` | CPU/GPU total vs allocated |\n| `ray_dashboard_status` | read | Ray | Ray dashboard | dashboard reachability/version |\n| `ray_job_list` | read | Ray | `GET /api/jobs/` | submitted jobs + status |\n| `gpu_utilization` | read | Ray | `GET /api/nodes` | per-node GPU count, utilisation %, memory (best-effort) |\n| `ray_job_cancel` | write (med) | Ray | `POST /api/jobs/{id}/stop` | cancel a running job |\n| `replica_restart` | write (**high**) | Ray | `GET/PUT /api/serve/applications/` | dry-run; restart a stuck replica |\n\n## Deploy lifecycle (4, write)\n\n| Tool | Risk | Backend | Endpoint | Undo / safety |\n|------|------|---------|----------|---------------|\n| `model_deploy` | med | Ray | `PUT /api/serve/applications/` | deploy an application |\n| `model_undeploy` | **high** | Ray | `DELETE/PUT /api/serve/applications/` | dry-run |\n| `deployment_redeploy` | **high** | Ray | `PUT /api/serve/applications/` | dry-run |\n| `routing_policy_update` | med | Ray | `PUT /api/serve/applications/` | reversible; prefix-aware / session-affinity routing to fix cache locality |\n\n## Cost (read, 1)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `cost_per_token` | vLLM | `GET /metrics` + GPU $/hr | deterministic $/1M tokens from measured throughput × GPU hourly rate |\n\n## SGLang / TGI writes (control-plane teaching error)\n\nSGLang and TGI are **single-process servers** with no Ray Serve control plane, so\nthe Ray-shaped write groups above (scale / drain / autoscale / deploy / redeploy /\nrouting / job-cancel / replica-restart) do not apply to them. Attempting one\nagainst a SGLang/TGI target raises `EngineCapabilityError` with a teaching message\npointing at a real horizontal-scale layer (Ray Serve / Kubernetes / a load\nbalancer). Their supported surface is the **engine-agnostic read** group above.\n\n## Out of scope (by design)\n\n- Cluster **provisioning** (spinning up GPU nodes, driver/CUDA install)\n- vLLM/Ray **install or version upgrades**\n- Multi-node **drain / reboot orchestration** (single-replica drain only, unverified at multi-node scale)\n- Non-inference infrastructure (use the appropriate other AIops-tools line)\n\nWant one of these? Open an issue or PR — feedback and contributions welcome.\n\nFile v0.10.2:references/cli-reference.md\n\n# inference-aiops CLI reference\n\n> Serving engines: vLLM (OpenAI API + Prometheus `/metrics`,\n> default 8000) with its Ray dashboard control plane (Serve + Jobs, default 8265),\n> plus single-process SGLang (default 30000) and TGI (default 8080); endpoints\n> need live verification.\n>\n> The CLI is a convenience subset. The full 35-tool surface — including the\n> engine-agnostic reads (`engine_health`, `engine_inventory`,\n> `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) that\n> cover SGLang/TGI — is via the MCP server (`inference-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\ninference-aiops init                      # interactive wizard: engine (vllm/sglang/tgi) + host + port\ninference-aiops doctor [--skip-auth]      # config + secret store + connectivity — vLLM: Ray + vLLM; SGLang/TGI: engine health + inventory\ninference-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.inference-aiops/secrets.enc — only if a token is used)\n\n```bash\ninference-aiops secret set <target> [--value <token>]  # store a bearer token (hidden prompt if no --value)\ninference-aiops secret list                            # names only — values never shown\ninference-aiops secret rm <target>\ninference-aiops secret migrate                         # import legacy plaintext env (INFERENCE_<T>_TOKEN)\ninference-aiops secret rotate-password                 # re-encrypt under a new master password\n```\n\n## Read commands\n\n```bash\ninference-aiops overview [--target <t>]        # Serve deployments + total replicas + queue backpressure\ninference-aiops serve list                     # Ray Serve deployments + replica counts\ninference-aiops serve status <application> <deployment>   # one deployment's status + replica count\ninference-aiops metrics requests               # TTFT / TPOT / e2e latency + token totals (from vLLM /metrics)\ninference-aiops metrics queue                  # running vs waiting requests (backpressure)\ninference-aiops metrics diagnose               # flagship RCA: ranked cause of a latency spike + the knob to turn\n```\n\n## Write commands (governed; risk tier in parentheses)\n\n```bash\ninference-aiops serve scale <application> <deployment> <num_replicas>   # (med) reversible\ninference-aiops serve scale-to-zero <application> <deployment> [--dry-run]   # (high) --dry-run + double confirm; strands ingress\n```\n\n> The remaining writes — `scale_replicas_down`, `drain_replica`,\n> `autoscale_config_update`, `lora_load` / `lora_unload`, `model_sleep` /\n> `model_wake` (dev-mode servers only),\n> `ray_job_cancel`, `replica_restart`, `model_deploy` / `model_undeploy`,\n> `deployment_redeploy`, `routing_policy_update` — are exposed via the MCP\n> server. High-risk ones support a dry-run preview.\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the default/first target)\n- `--dry-run` — print the API call that would be made, change nothing\n- State-changing commands (e.g. `serve scale-to-zero`) require two confirmations\n\nFile v0.10.2:references/setup-guide.md\n\n# inference-aiops setup & security guide\n\n> Not yet validated against a live cluster — see `docs/VERIFICATION.md`.\n\n## 1. Install\n\n```bash\nuv tool install inference-aiops\n```\n\n## 2. (Optional) create a bearer token\n\nA bearer token is **optional** — many vLLM / Ray stacks run open. Only create\none if your API requires it (e.g. vLLM started with `--api-key`, or an\nauthenticating proxy in front of the Ray dashboard). inference-aiops sends it as\n`Authorization: Bearer <token>` to both the Ray dashboard and vLLM.\n\n## 3. Onboard\n\n```bash\ninference-aiops init\n```\n\nThe wizard collects (non-secret) connection details into\n`~/.inference-aiops/config.yaml`: the **serving engine** (`vllm` / `sglang` /\n`tgi`), a **host**, the engine's **port**, and the **scheme** (http/https). For\nthe vLLM engine it also collects the **Ray dashboard port** (default 8265) — its\ncontrol plane. A bearer token is stored **encrypted** into\n`~/.inference-aiops/secrets.enc` **only if the API requires one**. Example config:\n\n```yaml\ntargets:\n  - name: prod                 # vLLM + Ray Serve control plane\n    host: 10.0.0.20\n    engine: vllm\n    scheme: http\n    ray_port: 8265\n    vllm_port: 8000            # vLLM engine port (legacy key; == engine_port)\n    verify_ssl: false          # self-signed lab certs only\n  - name: sg                    # SGLang (single-process, no Ray)\n    host: 10.0.0.21\n    engine: sglang\n    engine_port: 30000         # SGLang default\n  - name: edge                  # TGI (single-process, no Ray)\n    host: 10.0.0.22\n    engine: tgi\n    engine_port: 8080          # TGI default\n```\n\nDefault engine ports: **vLLM 8000, SGLang 30000, TGI 8080**. `engine_port` (or the\nlegacy `vllm_port`) sets the engine's HTTP port. Only the **vLLM** engine uses\n`ray_port`; SGLang and TGI are single-process servers with no Ray dashboard, so\ntheir scale/drain writes raise a teaching error rather than issuing an impossible\ncontrol-plane call.\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nIf a token is stored, export the master password so the encrypted store unlocks\nwithout a prompt (no token stored → nothing to export):\n\n```bash\nexport INFERENCE_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## 5. Laptop self-test (~80% of the tool, free)\n\nMost of the tool self-tests on a laptop with no cloud GPUs:\n\n- **vLLM** — run on a single GPU, or use a CPU-mock, exposing the OpenAI API and\n  Prometheus `/metrics` (default port 8000).\n- **Ray** — one local head node: `ray start --head` (Ray dashboard on 8265),\n  serving a small Serve app.\n\nPoint a target at `host: 127.0.0.1` with `ray_port: 8265` / `vllm_port: 8000`\nand run `inference-aiops doctor`. Reads, RCA, and most scaling ops exercise\nend-to-end. Still unverified: multi-GPU tensor/pipeline-parallel deployments,\nreal GPU thermal/throttle telemetry, and multi-node drain (see\n`docs/VERIFICATION.md`).\n\n## Credential security (when a token is used)\n\n- The token is **never** written to disk in plaintext. It lives only in\n  `~/.inference-aiops/secrets.enc`, encrypted with Fernet (AES-128-CBC + HMAC),\n  the key derived from your master password via scrypt. Only a per-store random\n  salt and the ciphertext are on disk (chmod 600); the master password is never\n  stored.\n- A legacy plaintext env var `INFERENCE_<TARGET_NAME_UPPER>_TOKEN` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `inference-aiops secret migrate`.\n- The token is held only in memory during a session and is never logged or\n  echoed; exception text and tracebacks are scrubbed of secret-shaped strings\n  before being written to the audit log.\n\n## Optional audit annotations\n\nThe tool does not require an approver — whether a high-risk write (scale-down,\nscale-to-zero, drain, LoRA unload, hot-swap, replica restart, undeploy, redeploy)\nshould happen is the agent's decision or the connecting environment's permission.\nIf you want the audit row to carry who ran a change and why, set these; they are\nrecorded when present and never required:\n\n```bash\nexport INFERENCE_AUDIT_APPROVED_BY='you'\nexport INFERENCE_AUDIT_RATIONALE='off-peak cost save'\n```\n\n## Governance harness state\n\nState lives under `~/.inference-aiops/` (relocate with `INFERENCE_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with the risk tier (a descriptive\n  label, not a gate) and any approver/rationale annotation\n- `undo.db` — inverse descriptors for reversible writes (scale, autoscale-config,\n  routing, hot-swap, LoRA load)\n- budget / runaway guard — a safety backstop (not authorization): caps cumulative\n  tool calls and wall-time; trips on tight poll/retry loops\n\n## Verify\n\n```bash\ninference-aiops doctor\n```\n\n`doctor` checks the config file, the encrypted store and its permissions (if a\ntoken is configured), and — unless `--skip-auth` — connectivity. For a **vLLM**\ntarget it probes the **Ray dashboard** and **vLLM** independently, so a half-up\ncluster is reported precisely; for a **SGLang / TGI** target it probes the\nengine's health endpoint and running-model inventory (no Ray).\n\nFile v0.10.2:skill-card.md\n\n## Description:\n\nInference AIops helps agents operate GPU inference clusters across vLLM, Ray Serve, SGLang, and TGI by inspecting health and metrics, diagnosing latency or utilization issues, and guiding governed operational changes such as scaling, draining, LoRA operations, and deployment lifecycle actions.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers and inference platform operators use this skill to monitor GPU inference services, diagnose latency or utilization problems, and perform controlled operational actions on vLLM/Ray Serve deployments or read-only checks for SGLang and TGI services.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Disruptive write controls can scale, drain, restart, unload, undeploy, or otherwise affect live inference service without an enforced read-only authorization gate.\n\nMitigation: Use the skill only against clusters where the agent is authorized to make operational changes, and restrict network/API access when an observe-only posture is required.\n\nRisk: The release installs an external executable that security evidence describes as unpinned.\n\nMitigation: Pin the inference-aiops package to version 0.10.2 or another approved artifact before deployment.\n\nRisk: Credentials and serving control-plane access can affect both Ray and inference endpoints.\n\nMitigation: Use separate least-privilege credentials, keep HTTPS certificate verification enabled where possible, and avoid plaintext token environment fallbacks.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/inference-aiops)\n- [Project homepage](https://github.com/AIops-tools/Inference-AIops)\n- [Capabilities reference](references/capabilities.md)\n- [CLI reference](references/cli-reference.md)\n- [Setup and security guide](references/setup-guide.md)\n- [Agent guardrails](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Analysis, Shell commands, Configuration]\n\n**Output Format:** [Markdown with inline shell commands, operational recommendations, and structured observations from tool results]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include dry-run recommendations, audit guidance, risk-tier notes, and follow-up verification steps for inference operations.]\n\n## Skill Version(s):\n\n0.10.2 (source: ClawHub release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.1: 7 files, 19747 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (3111b), SKILL.md (17216b), _meta.json (135b)\n\nFile v0.10.1:SKILL.md\n\n---\nname: inference-aiops\nslug: inference-aiops\ndisplayName: \"Inference AIops\"\nsummary: \"Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Inference-AIops\ntags: [aiops, mcp, governance, inference]\ndescription: >\n  Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens.\n  Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster.\n  Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray).\n  Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).\ninstaller:\n  kind: uv\n  package: inference-aiops\nargument-hint: \"[deployment/model name or describe your inference-cluster task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"inference-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"INFERENCE_AIOPS_CONFIG\",\"INFERENCE_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Inference-AIops\",\"emoji\":\"🚀\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.inference-aiops/ (relocatable via INFERENCE_AIOPS_HOME).\n  Auth: a bearer token is OPTIONAL — many vLLM / Ray stacks run open. When the API requires one it is stored ENCRYPTED in ~/.inference-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'inference-aiops init' to onboard, or 'inference-aiops secret set <target>' to add one. The store is unlocked by a master password from INFERENCE_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var INFERENCE_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback (migrate with 'inference-aiops secret migrate'). The token is sent as an Authorization: Bearer header at request time and held only in memory; it is never logged or echoed.\n  State-changing operations require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + risk-tier label — it records, not authorizes). The fragile prod ops — scale_replicas_down, scale_to_zero, drain_replica, lora_unload, model_sleep, replica_restart, model_undeploy, deployment_redeploy — are high-risk with a dry_run preview; reversible writes (scale, autoscale-config, routing, sleep, LoRA load) record an undo descriptor.\n  Engines: vLLM (with its Ray Serve control plane), SGLang, and TGI. SGLang/TGI are single-process servers with engine-agnostic observability (health, running-model inventory, request metrics, queue depth, latency RCA); Ray-shaped scale/drain writes are vLLM-only and raise a teaching error on a SGLang/TGI target.\n  Metrics: each engine's Prometheus /metrics endpoint is parsed directly — no Prometheus server is required.\n  Webhooks: none — no outbound calls beyond the configured Ray dashboard and vLLM services.\n  SSL: verify_ssl defaults to true; disable only for self-signed lab certificates.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked vLLM/Ray responses; unverified against multi-GPU tensor/pipeline-parallel deployments, real GPU thermal/throttle telemetry, and multi-node drain (see docs/VERIFICATION.md).\n---\n\n# Inference AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor.** Product and trademark names belong to their owners. Source at [github.com/AIops-tools/Inference-AIops](https://github.com/AIops-tools/Inference-AIops) under the MIT license.\n\nGoverned GPU-inference operations for **vLLM** (OpenAI API + Prometheus `/metrics`) and **Ray Serve / Ray Jobs** (Ray dashboard), plus the single-process serving engines **SGLang** and **TGI** — **39 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.inference-aiops/`, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels on every audit row. The flagship `diagnose_latency_spike` folds queue depth + KV-cache pressure + prefix-cache locality into a ranked cause and the specific knob to turn; the engine-agnostic `diagnose_engine_latency` does the same across whatever signals SGLang/TGI expose. Each engine's Prometheus `/metrics` is parsed directly — **no Prometheus server required**.\n\n> **Standalone**: the governance harness is bundled in the package (`inference_aiops.governance`) — no external skill-family dependency. A bearer token is **optional** (many stacks run open).\n\n## What This Skill Does\n\n| Group | Tools | Count | Read or Write |\n|-------|-------|:-----:|:-------------:|\n| **Metrics & RCA** (vLLM) | request metrics, queue depth, KV-cache stats, diagnose latency spike, diagnose low utilisation | 5 | 5 read |\n| **Engine-agnostic** (vLLM/SGLang/TGI) | engine health, engine inventory, engine request metrics, engine queue depth, diagnose engine latency | 5 | 5 read |\n| **Ray Serve (read)** | deployment list, deployment status, replica list, autoscale config get | 4 | 4 read |\n| **Ray Serve (write)** | scale up (med), scale down (high), scale-to-zero (high), autoscale config update (med), drain replica (high) | 5 | 5 write |\n| **Models / vLLM** | model list, model info, LoRA load (med), LoRA unload (high), base hot-swap (high) | 5 | 2 read / 3 write |\n| **Ray cluster / jobs / GPU** | cluster resources, dashboard status, job list, GPU utilisation, job cancel (med), replica restart (high) | 6 | 4 read / 2 write |\n| **Deploy lifecycle** | deploy (med), undeploy (high), redeploy (high), routing policy update (med) | 4 | 4 write |\n| **Cost** | cost per token | 1 | 1 read |\n\n**23 read, 16 write**, plus `undo_list` / `undo_apply` — **39 MCP tools** in total. The high-risk writes support `dry_run` + double-confirm; reversible writes record an undo descriptor. The engine-agnostic reads cover any engine; the Ray Serve / cluster / deploy write groups are vLLM-only and teach-and-refuse on a SGLang/TGI target (single-process engines have no Ray control plane).\n\n## Quick Install\n\n```bash\nuv tool install inference-aiops\ninference-aiops init       # interactive wizard: engine (vllm/sglang/tgi) + host + port + scheme (token optional)\ninference-aiops doctor     # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@aiops-tools/inference-aiops\nopenclaw skills info inference-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Triage a cluster (`overview`): Serve deployments, total replicas, queue backpressure\n- Diagnose slow inference (`metrics diagnose` / `diagnose_latency_spike`): rank the cause (queue depth vs KV-cache preemption vs prefix-cache locality) and get the knob to turn\n- Find idle GPUs and over-provisioned replicas (`diagnose_low_utilization`)\n- Scale a Ray Serve deployment up/down, **scale-to-zero** to stop cost bleed, or update autoscale bounds\n- **Drain** a replica gracefully before a node reboot (finishes in-flight requests)\n- Load/unload a **LoRA** adapter; **hot-swap** a base model (Sleep-Mode swap, captures the prior model)\n- Inspect GPU utilisation per node, list/cancel Ray jobs, restart a stuck replica\n- Compute **cost per million tokens** from throughput × GPU $/hr\n- Observe an **SGLang** or **TGI** server (`engine_health`, `engine_inventory`, `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) — single-process engines with no Ray control plane\n\n**Do NOT use for** non-inference infrastructure (hypervisors, storage appliances, backup products, general container workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools. This skill is scoped to GPU inference serving (vLLM + Ray).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| vLLM / Ray Serve inference: latency RCA, autoscale, drain, LoRA, cost/token | **inference-aiops** (this skill) |\n| SGLang / TGI serving: health, running-model inventory, request metrics, queue depth, latency RCA | **inference-aiops** (this skill — engine-agnostic reads) |\n| Any non-inference infrastructure (hypervisor, storage, backup, general clusters, network, OT) | the appropriate **other AIops-tools** line |\n\n## Common Workflows\n\n### 1. \"Inference got slow this afternoon\" (flagship RCA → the right knob)\n\n1. `inference-aiops doctor` → confirm the vLLM endpoint and Ray dashboard are actually reachable before blaming the model\n2. `inference-aiops overview` → Serve deployments, total replicas, and whether queue backpressure is cluster-wide or one deployment\n3. `inference-aiops metrics diagnose` (MCP: `diagnose_latency_spike`) → a **ranked** cause with the measured numbers: is `waiting` queue depth high (backpressure)? Are there KV-cache **preemptions** (`kv_cache_stats`)? Has the **prefix-cache hit rate** dropped (routing lost locality)?\n4. Turn the knob the RCA names, not a guess:\n   - backpressure → `inference-aiops serve scale <app> <deployment> --replicas N` (`scale_replicas_up`, reversible, prior count captured)\n   - KV-cache preemption → `autoscale_config_update` to lower the concurrent-request cap (reversible, prior config captured)\n   - lost locality → `routing_policy_update` to prefix-aware / session-affinity (reversible)\n5. Re-check `inference-aiops metrics requests` (TTFT / TPOT / e2e) and `inference-aiops metrics queue` to confirm the p99 actually moved\n6. **Failure branch**: if the fix makes it worse, `inference-aiops undo list` → `inference-aiops undo apply <id>` restores the exact prior replica count\n\nArchive v0.10.0: 7 files, 19381 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (2639b), SKILL.md (16898b), _meta.json (135b)\n\nArchive v0.9.0: 7 files, 19425 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (2711b), SKILL.md (17023b), _meta.json (134b)\n\nArchive v0.8.0: 7 files, 19545 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (3004b), SKILL.md (17023b), _meta.json (134b)\n\nArchive v0.7.0: 7 files, 19678 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (3279b), SKILL.md (17023b), _meta.json (134b)\n\nArchive v0.6.0: 7 files, 19582 bytes\n\nFiles: references/agent-guardrails.md (7025b), references/capabilities.md (7615b), references/cli-reference.md (3066b), references/setup-guide.md (5071b), skill-card.md (2977b), SKILL.md (17023b), _meta.json (134b)\n\nArchive v0.5.0: 7 files, 18786 bytes\n\nFiles: references/agent-guardrails.md (6449b), references/capabilities.md (7615b), references/cli-reference.md (3090b), references/setup-guide.md (4825b), skill-card.md (2708b), SKILL.md (16391b), _meta.json (134b)","readmeExcerpt":"Skill: inference-aiops Owner: zw008 Summary: Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token t","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install inference-aiops\ninference-aiops init       # interactive wizard: engine (vllm/sglang/tgi) + host + port + scheme (token optional)\ninference-aiops doctor     # vLLM: probes Ray + vLLM; SGLang/TGI: engine health + inventory"},{"language":"bash","snippet":"openclaw plugins install clawhub:@zw008/inference-aiops\nopenclaw skills info inference-aiops          # expect: Visible to model: yes"},{"language":"text","snippet":"You operate a GPU inference cluster through the inference-aiops MCP tools\n(vLLM / SGLang / TGI serving engines, plus a Ray Serve control plane).\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a tool.\n  Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"healthy\".\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null metric means the engine does not expose that signal. Report it as \"not\n  available\" — never substitute 0, and never compare a null against a threshold.\n- Report values exactly as returned. Do not normalise or prettify model ids,\n  deployment names, replica states, or Ray job statuses.\n- When diagnose_engine_latency or diagnose_latency_spike returns probableCauses,\n  work through them in the order given and cite the measured number in each\n  cause's \"signal\" — do not substitute your own theory of the bottleneck.\n\n- Only `scale_to_zero` has a CLI command; every other traffic-affecting tool is MCP-only\n  and nothing will ask you to confirm it. Call it with `dry_run=True` first, show the\n  operator what would change, and wait for an explicit go-ahead.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a latency, throughput, or capacity problem unless a tool result\n  supports it. High GPU utilisation is not by itself a fault.\n- Do not confuse the identifier kinds: a Ray *application* name, a *deployment*\n  name within it, a *replica* id, "},{"language":"bash","snippet":"inference-aiops doctor"},{"language":"bash","snippet":"export INFERENCE_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport INFERENCE_AUDIT_RATIONALE=\"scaling llm-app down for the maintenance window\""},{"language":"bash","snippet":"inference-aiops init                      # interactive wizard: engine (vllm/sglang/tgi) + host + port\ninference-aiops doctor [--skip-auth]      # config + secret store + connectivity — vLLM: Ray + vLLM; SGLang/TGI: engine health + inventory\ninference-aiops mcp                       # start the MCP server (stdio transport)"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: inference-aiops\nslug: inference-aiops\ndisplayName: \"Inference AIops\"\nsummary: \"Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Inference-AIops\ntags: [aiops, mcp, governance, inference]\ndescription: >\n  Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens.\n  Always use this skill for \"why is inference slow\", \"TTFT spike\", \"latency spike\", \"GPU underutilised\", \"scale down the deployment\", \"scale to zero\", \"drain a replica before a reboot\", \"hot-swap the base model\", \"load a LoRA adapter\", \"KV cache pressure\", \"prefix cache hit rate\", \"queue backpressure\", \"autoscale config\", \"SGLang health\", \"TGI metrics\", or \"cost per token\" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster.\n  Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray).\n  Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).\ninstaller:\n  kind: uv\n  package: inference-aiops\nargument-hint: \"[deployment/model name or describe your inference-cluster task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"inference-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"INFERENCE_AIOPS_CONFIG\",\"INFERENCE_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Inference-AIops\",\"emoji\":\"🚀\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.inference-aiops/ (relocatable via INFERENCE_AIOPS_HOME).\n  Auth: a bearer token is OPTIONAL — many vLLM / Ray stacks run open. When the API requires one it is stored ENCRYPTED in ~/.inference-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'i"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"inference-aiops\",\n  \"version\": \"0.10.4\",\n  \"publishedAt\": 1789601150857\n}"},{"path":"references/agent-guardrails.md","content":"# Agent guardrails — running inference-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the environment's. The tool\ndoes not gate it — there is no read-only switch and no approval prompt to\nconfigure. The two right places to control read vs write:\n\n- **The environment you connect it to.** Restrict the network path so the tool\n  can only reach the read/metrics endpoints, or run the Ray dashboard without its\n  job-submission API. A write then fails at the server, which is the only place\n  the permission actually lives — no skill-side flag can be argued around by a\n  model, but a blocked endpoint cannot.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the engine or Ray dashboard did not return comes back as `null`, never as `\"\"`. An absent job `entrypoint`, a model's `parent` adapter, a replica `state`, or a server-info `version` is distinguishable from an empty one. |\n| \"Tell me if the output was cut off\" | `ray_job_list` returns `{\"jobs\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`. Truncation is measured against the full fetch, not guessed from a length coincidence. |\n| \"Say when a metric isn't available\" | Signals the engine does not expose come back as `null` rather than `0`. SGLang and TGI expose fewer metrics than vLLM; `diagnose_engine_latency` skips a signal it cannot read instead of fabricating it, and `signalsChecked` shows exactly what it looked at. |\n| \"Don't suggest scaling on an engine that can't scale\" | Multi-replica scale / drain / autoscale are Ray Serve control-plane actions. On a single-process engine (SGLang, TGI) those tools raise `EngineCapabilityError` with an explanation, rather than issuing a call that could never succeed. |\n| \"Confirm before anything disruptive\" | Every traffic-affecting operation (`model_undeploy`, `deployment_redeploy`, `scale_to_zero`, `scale_replicas_down`, `drain_replica`, `replica_restart`, `lora_unload`, `model_sleep`) takes `dry_run=True` for a preview and is `risk=high`. ⚠️ **The double confirm"},{"path":"references/capabilities.md","content":"# inference-aiops capabilities\n\n> 39 MCP tools (23 read, 16 write, 2 undo). Serving engines:\n> **vLLM** (OpenAI API + Prometheus `/metrics`, default 8000) with its **Ray**\n> dashboard control plane (Serve + Jobs, default 8265), plus the single-process\n> engines **SGLang** (OpenAI API + `/get_server_info` + Prometheus `/metrics`,\n> default 30000) and **TGI** (`/info` + Prometheus `/metrics`, default 8080).\n> Endpoints modelled against those APIs; need live verification.\n\n## Metrics & RCA — vLLM (read, 5)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `request_metrics` | vLLM | `GET /metrics` | TTFT, TPOT, e2e latency (avg/p50/p90/p99), prompt/generation token totals, request counts |\n| `queue_depth` | vLLM | `GET /metrics` | running vs waiting requests (backpressure), scheduler state |\n| `kv_cache_stats` | vLLM | `GET /metrics` | KV-cache utilisation %, prefix-cache hit rate, preemption count |\n| `diagnose_latency_spike` | vLLM | `GET /metrics` (fold) | **ranked cause** (queue backpressure / KV-cache preemption / prefix-cache locality) + the specific knob to turn |\n| `diagnose_low_utilization` | vLLM | `GET /metrics` (fold) | idle-GPU / over-provisioned / routing-stranded diagnosis + what to scale down |\n\n## Engine-agnostic — vLLM / SGLang / TGI (read, 5)\n\nWork against **any** supported engine, reading each engine's own paths and metric\nnames (vLLM `vllm:*`, SGLang `sglang:*`, TGI `tgi_*`). A signal an engine does not\nexpose (e.g. TGI has no TTFT or KV-cache metric) degrades to `null` rather than\nbeing guessed.\n\n| Tool | Endpoint(s) | Returns |\n|------|-------------|---------|\n| `engine_health` | `GET /health` | engine liveness (`healthy` bool) + engine label |\n| `engine_inventory` | `GET /v1/models` (vLLM/SGLang) or `/info` (TGI); `/get_server_info` (SGLang) | running-model id(s) + best-effort server info (model, version, max concurrency) |\n| `engine_request_metrics` | `GET /metrics` | TTFT / TPOT / e2e latency + generation-token totals, per engine's exposition (null where unexposed) |\n| `engine_queue_depth` | `GET /metrics` | running vs waiting requests + backpressure flag (SGLang `num_queue_reqs`, TGI `tgi_queue_size`) |\n| `diagnose_engine_latency` | `GET /metrics` (fold) | **ranked cause** across the signals the engine exposes (queue backpressure / KV-token-cache pressure / cache locality) + the knob to turn |\n\n## Ray Serve — read (4)\n\n| Tool | Backend | Endpoint | Returns |\n|------|---------|----------|---------|\n| `serve_deployment_list` | Ray | `GET /api/serve/applications/` | all Serve deployments: status, replica count, target |\n| `deployment_status` | Ray | `GET /api/serve/applications/` | one deployment's status + current/target replica count |\n| `replica_list` | Ray | `GET /api/serve/applications/` | per-replica id, state, node |\n| `autoscale_config_get` | Ray | `GET /api/serve/applications/` | min/max replicas, target ongoing requests |\n\n## Ray Serve — write (5)\n\n| Tool | Risk | Backend | Endpoint |"},{"path":"references/cli-reference.md","content":"# inference-aiops CLI reference\n\n> Serving engines: vLLM (OpenAI API + Prometheus `/metrics`,\n> default 8000) with its Ray dashboard control plane (Serve + Jobs, default 8265),\n> plus single-process SGLang (default 30000) and TGI (default 8080); endpoints\n> need live verification.\n>\n> The CLI is a convenience subset. The full 35-tool surface — including the\n> engine-agnostic reads (`engine_health`, `engine_inventory`,\n> `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) that\n> cover SGLang/TGI — is via the MCP server (`inference-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\ninference-aiops init                      # interactive wizard: engine (vllm/sglang/tgi) + host + port\ninference-aiops doctor [--skip-auth]      # config + secret store + connectivity — vLLM: Ray + vLLM; SGLang/TGI: engine health + inventory\ninference-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.inference-aiops/secrets.enc — only if a token is used)\n\n```bash\ninference-aiops secret set <target> [--value <token>]  # store a bearer token (hidden prompt if no --value)\ninference-aiops secret list                            # names only — values never shown\ninference-aiops secret rm <target>\ninference-aiops secret migrate                         # import legacy plaintext env (INFERENCE_<T>_TOKEN)\ninference-aiops secret rotate-password                 # re-encrypt under a new master password\n```\n\n## Read commands\n\n```bash\ninference-aiops overview [--target <t>]        # Serve deployments + total replicas + queue backpressure\ninference-aiops serve list                     # Ray Serve deployments + replica counts\ninference-aiops serve status <application> <deployment>   # one deployment's status + replica count\ninference-aiops metrics requests               # TTFT / TPOT / e2e latency + token totals (from vLLM /metrics)\ninference-aiops metrics queue                  # running vs waiting requests (backpressure)\ninference-aiops metrics diagnose               # flagship RCA: ranked cause of a latency spike + the knob to turn\n```\n\n## Write commands (governed; risk tier in parentheses)\n\n```bash\ninference-aiops serve scale <application> <deployment> <num_replicas>   # (med) reversible\ninference-aiops serve scale-to-zero <application> <deployment> [--dry-run]   # (high) --dry-run + double confirm; strands ingress\n```\n\n> The remaining writes — `scale_replicas_down`, `drain_replica`,\n> `autoscale_config_update`, `lora_load` / `lora_unload`, `model_sleep` /\n> `model_wake` (dev-mode servers only),\n> `ray_job_cancel`, `replica_restart`, `model_deploy` / `model_undeploy`,\n> `deployment_redeploy`, `routing_policy_update` — are exposed via the MCP\n> server. High-risk ones support a dry-run preview.\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the default/first target)\n- `--dry-run` — print the API call that would be made, change nothing\n- State-changing commands (e.g. `"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2253,"uniquenessScore":37,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T07:16:11.212Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T10:44:06.314Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}