inference-aiops
Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens. Always use this skill for "why is inference slow", "TTFT spike", "latency spike", "GPU underutilised", "scale down the deployment", "scale to zero", "drain a replica before a reboot", "hot-swap the base model", "load a LoRA adapter", "KV cache pressure", "prefix cache hit rate", "queue backpressure", "autoscale config", "SGLang health", "TGI metrics", or "cost per token" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster. Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray). Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).
Rank
62
Safety
84
Downloads
1.6k
Updated
Oct 10, 2026
Version
0.10.4
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.6K downloadsadoption · observed Oct 10, 2026
- Latest release
- 0.10.4release · observed Sep 16, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:inference-aiops- Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:inference-aiops` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/inference-aiops before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/snapshot"
Documentation
CLAWHUB
150,199 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: inference-aiops
slug: inference-aiops
displayName: "Inference AIops"
summary: "Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 39 tools."
license: MIT
homepage: https://github.com/AIops-tools/Inference-AIops
tags: [aiops, mcp, governance, inference]
description: >
Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats (utilisation, prefix-cache hit rate, preemptions), the flagship latency root-cause analysis (diagnose_latency_spike / diagnose_engine_latency) and low-utilisation RCA, engine-agnostic health + running-model inventory across vLLM/SGLang/TGI, Ray Serve autoscaling and scaling (scale up/down, scale-to-zero, drain a replica), LoRA load/unload, base-model hot-swap, deploy/undeploy/redeploy, prefix-aware routing, GPU utilisation, Ray jobs, and cost per million tokens.
Always use this skill for "why is inference slow", "TTFT spike", "latency spike", "GPU underutilised", "scale down the deployment", "scale to zero", "drain a replica before a reboot", "hot-swap the base model", "load a LoRA adapter", "KV cache pressure", "prefix cache hit rate", "queue backpressure", "autoscale config", "SGLang health", "TGI metrics", or "cost per token" when the context is a vLLM / SGLang / TGI / Ray Serve inference cluster.
Do NOT use for non-inference infrastructure (hypervisors, storage appliances, backup products, general container/cluster workloads, network devices, or OT/industrial equipment) — those belong to other AIops-tools; this skill is scoped to GPU inference serving (vLLM + Ray).
Governed vLLM + Ray inference operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).
installer:
kind: uv
package: inference-aiops
argument-hint: "[deployment/model name or describe your inference-cluster task]"
allowed-tools:
- Bash
metadata: {"openclaw":{"requires":{"anyBins":["inference-aiops","uvx"]},"optional":{"env":["INFERENCE_AIOPS_CONFIG","INFERENCE_AIOPS_MASTER_PASSWORD"]},"homepage":"https://github.com/AIops-tools/Inference-AIops","emoji":"🚀","os":["macos","linux"]}}
compatibility: >
Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.
All write operations are audited to a local SQLite DB under ~/.inference-aiops/ (relocatable via INFERENCE_AIOPS_HOME).
Auth: a bearer token is OPTIONAL — many vLLM / Ray stacks run open. When the API requires one it is stored ENCRYPTED in ~/.inference-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'i_meta.json
{
"ownerId": "kn7b067awq2s97bn3d7p5qfhw5827pxc",
"slug": "inference-aiops",
"version": "0.10.4",
"publishedAt": 1789601150857
}references/agent-guardrails.md
# Agent guardrails — running inference-aiops with a smaller / local model
If you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,
Ollama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably
better results with a short system prompt. This page gives you one, and — more
importantly — tells you which guardrails you **no longer need to write**, because
the tool now enforces them itself.
The distinction matters. A guardrail in a prompt is a request. A guardrail in the
harness is a guarantee. Anything below that we could move into the harness, we did.
## Authorization is not this tool's job — decide it where it belongs
Whether a write should happen is your decision, or the environment's. The tool
does not gate it — there is no read-only switch and no approval prompt to
configure. The two right places to control read vs write:
- **The environment you connect it to.** Restrict the network path so the tool
can only reach the read/metrics endpoints, or run the Ray dashboard without its
job-submission API. A write then fails at the server, which is the only place
the permission actually lives — no skill-side flag can be argued around by a
model, but a blocked endpoint cannot.
- **Your agent's system prompt.** If you want an observe-only session, tell the
model not to call the write tools (they are clearly tagged `[WRITE]`).
What the tool *does* guarantee is that you can always see what happened:
## What the tool now enforces — do not waste prompt budget on these
| You might be tempted to prompt | Why you don't need to |
|---|---|
| "Don't invent a value when a field is missing" | A field the engine or Ray dashboard did not return comes back as `null`, never as `""`. An absent job `entrypoint`, a model's `parent` adapter, a replica `state`, or a server-info `version` is distinguishable from an empty one. |
| "Tell me if the output was cut off" | `ray_job_list` returns `{"jobs": [...], "returned": N, "limit": L, "truncated": true/false}`. Truncation is measured against the full fetch, not guessed from a length coincidence. |
| "Say when a metric isn't available" | Signals the engine does not expose come back as `null` rather than `0`. SGLang and TGI expose fewer metrics than vLLM; `diagnose_engine_latency` skips a signal it cannot read instead of fabricating it, and `signalsChecked` shows exactly what it looked at. |
| "Don't suggest scaling on an engine that can't scale" | Multi-replica scale / drain / autoscale are Ray Serve control-plane actions. On a single-process engine (SGLang, TGI) those tools raise `EngineCapabilityError` with an explanation, rather than issuing a call that could never succeed. |
| "Confirm before anything disruptive" | Every traffic-affecting operation (`model_undeploy`, `deployment_redeploy`, `scale_to_zero`, `scale_replicas_down`, `drain_replica`, `replica_restart`, `lora_unload`, `model_sleep`) takes `dry_run=True` for a preview and is `risk=high`. ⚠️ **The double confirmreferences/capabilities.md
# inference-aiops capabilities > 39 MCP tools (23 read, 16 write, 2 undo). Serving engines: > **vLLM** (OpenAI API + Prometheus `/metrics`, default 8000) with its **Ray** > dashboard control plane (Serve + Jobs, default 8265), plus the single-process > engines **SGLang** (OpenAI API + `/get_server_info` + Prometheus `/metrics`, > default 30000) and **TGI** (`/info` + Prometheus `/metrics`, default 8080). > Endpoints modelled against those APIs; need live verification. ## Metrics & RCA — vLLM (read, 5) | Tool | Backend | Endpoint | Returns | |------|---------|----------|---------| | `request_metrics` | vLLM | `GET /metrics` | TTFT, TPOT, e2e latency (avg/p50/p90/p99), prompt/generation token totals, request counts | | `queue_depth` | vLLM | `GET /metrics` | running vs waiting requests (backpressure), scheduler state | | `kv_cache_stats` | vLLM | `GET /metrics` | KV-cache utilisation %, prefix-cache hit rate, preemption count | | `diagnose_latency_spike` | vLLM | `GET /metrics` (fold) | **ranked cause** (queue backpressure / KV-cache preemption / prefix-cache locality) + the specific knob to turn | | `diagnose_low_utilization` | vLLM | `GET /metrics` (fold) | idle-GPU / over-provisioned / routing-stranded diagnosis + what to scale down | ## Engine-agnostic — vLLM / SGLang / TGI (read, 5) Work against **any** supported engine, reading each engine's own paths and metric names (vLLM `vllm:*`, SGLang `sglang:*`, TGI `tgi_*`). A signal an engine does not expose (e.g. TGI has no TTFT or KV-cache metric) degrades to `null` rather than being guessed. | Tool | Endpoint(s) | Returns | |------|-------------|---------| | `engine_health` | `GET /health` | engine liveness (`healthy` bool) + engine label | | `engine_inventory` | `GET /v1/models` (vLLM/SGLang) or `/info` (TGI); `/get_server_info` (SGLang) | running-model id(s) + best-effort server info (model, version, max concurrency) | | `engine_request_metrics` | `GET /metrics` | TTFT / TPOT / e2e latency + generation-token totals, per engine's exposition (null where unexposed) | | `engine_queue_depth` | `GET /metrics` | running vs waiting requests + backpressure flag (SGLang `num_queue_reqs`, TGI `tgi_queue_size`) | | `diagnose_engine_latency` | `GET /metrics` (fold) | **ranked cause** across the signals the engine exposes (queue backpressure / KV-token-cache pressure / cache locality) + the knob to turn | ## Ray Serve — read (4) | Tool | Backend | Endpoint | Returns | |------|---------|----------|---------| | `serve_deployment_list` | Ray | `GET /api/serve/applications/` | all Serve deployments: status, replica count, target | | `deployment_status` | Ray | `GET /api/serve/applications/` | one deployment's status + current/target replica count | | `replica_list` | Ray | `GET /api/serve/applications/` | per-replica id, state, node | | `autoscale_config_get` | Ray | `GET /api/serve/applications/` | min/max replicas, target ongoing requests | ## Ray Serve — write (5) | Tool | Risk | Backend | Endpoint |
references/cli-reference.md
# inference-aiops CLI reference > Serving engines: vLLM (OpenAI API + Prometheus `/metrics`, > default 8000) with its Ray dashboard control plane (Serve + Jobs, default 8265), > plus single-process SGLang (default 30000) and TGI (default 8080); endpoints > need live verification. > > The CLI is a convenience subset. The full 35-tool surface — including the > engine-agnostic reads (`engine_health`, `engine_inventory`, > `engine_request_metrics`, `engine_queue_depth`, `diagnose_engine_latency`) that > cover SGLang/TGI — is via the MCP server (`inference-aiops mcp`). ## Setup & diagnostics ```bash inference-aiops init # interactive wizard: engine (vllm/sglang/tgi) + host + port inference-aiops doctor [--skip-auth] # config + secret store + connectivity — vLLM: Ray + vLLM; SGLang/TGI: engine health + inventory inference-aiops mcp # start the MCP server (stdio transport) ``` ## Secrets (encrypted store ~/.inference-aiops/secrets.enc — only if a token is used) ```bash inference-aiops secret set <target> [--value <token>] # store a bearer token (hidden prompt if no --value) inference-aiops secret list # names only — values never shown inference-aiops secret rm <target> inference-aiops secret migrate # import legacy plaintext env (INFERENCE_<T>_TOKEN) inference-aiops secret rotate-password # re-encrypt under a new master password ``` ## Read commands ```bash inference-aiops overview [--target <t>] # Serve deployments + total replicas + queue backpressure inference-aiops serve list # Ray Serve deployments + replica counts inference-aiops serve status <application> <deployment> # one deployment's status + replica count inference-aiops metrics requests # TTFT / TPOT / e2e latency + token totals (from vLLM /metrics) inference-aiops metrics queue # running vs waiting requests (backpressure) inference-aiops metrics diagnose # flagship RCA: ranked cause of a latency spike + the knob to turn ``` ## Write commands (governed; risk tier in parentheses) ```bash inference-aiops serve scale <application> <deployment> <num_replicas> # (med) reversible inference-aiops serve scale-to-zero <application> <deployment> [--dry-run] # (high) --dry-run + double confirm; strands ingress ``` > The remaining writes — `scale_replicas_down`, `drain_replica`, > `autoscale_config_update`, `lora_load` / `lora_unload`, `model_sleep` / > `model_wake` (dev-mode servers only), > `ray_job_cancel`, `replica_restart`, `model_deploy` / `model_undeploy`, > `deployment_redeploy`, `routing_policy_update` — are exposed via the MCP > server. High-risk ones support a dry-run preview. ## Common options - `--target, -t <name>` — target name from `config.yaml` (omit to use the default/first target) - `--dry-run` — print the API call that would be made, change nothing - State-changing commands (e.g. `
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/zw008/skills/inference-aiops",
"sourceUrl": "https://clawhub.ai/zw008/skills/inference-aiops",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T07:16:11.212Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T07:16:11.212Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.6K downloads",
"href": "https://clawhub.ai/zw008/inference-aiops",
"sourceUrl": "https://clawhub.ai/zw008/inference-aiops",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T07:16:11.212Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.10.4",
"href": "https://clawhub.ai/zw008/inference-aiops",
"sourceUrl": "https://clawhub.ai/zw008/inference-aiops",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-16T23:25:50.857Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-inference-aiops/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.10.4",
"description": "- Updated documentation in references/agent-guardrails.md. - Removed the skill-card.md file. - No functional changes to skill logic; documentation and metadata only.",
"href": "https://clawhub.ai/zw008/inference-aiops",
"sourceUrl": "https://clawhub.ai/zw008/inference-aiops",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-16T23:25:50.857Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
