ceph-aiops
Use this skill whenever the user needs to operate or diagnose a Ceph cluster via its ceph-mgr Dashboard REST API — decode a HEALTH_WARN/ERR state into cause + action (cluster_health), read the cluster status, inspect OSDs (tree/df/perf), placement groups (summary/stuck/scrub), pools (list/usable capacity), RBD images and snapshots, CephFS/MDS and RGW status, monitors/managers, slow ops and capacity forecast — plus governed writes (set cluster flags, reweight/mark-in/mark-out/purge OSDs, trigger scrubs, set pool quota/pg_num/autoscale/size, create/delete pools, create/delete RBD images and snapshots, throttle recovery/backfill). Always use this skill for "ceph health", "what does this HEALTH_WARN mean", "PG_DEGRADED / OSD_NEARFULL / SLOW_OPS / MON_DOWN", "ceph -s", "which OSD is most full", "drain an OSD", "purge an OSD", "stuck PGs", "overdue scrub", "pool usable capacity", "set pool size / quota", "rebalance is too slow / throttle backfill", "RBD image or snapshot", "MDS behind on trimming", "RGW large omap", "mon quorum", or "days to nearfull" when the context is a Ceph cluster (cephadm, hypervisor-bundled Ceph, or MicroCeph). Do NOT use when the target is not Ceph — a hypervisor, a different storage appliance, a backup product, a Kubernetes cluster, or a network device. Route those to the appropriate other AIops-tools skill (negative routing hint only). Common Ceph ops with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).
Rank
62
Safety
84
Downloads
2.0k
Updated
Oct 9, 2026
Version
0.11.5
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.11.5release · observed Sep 16, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:ceph-aiops- Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:ceph-aiops` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/ceph-aiops before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-zw008-ceph-aiops/snapshot"
Documentation
CLAWHUB
149,870 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: ceph-aiops
slug: ceph-aiops
displayName: "Ceph AIops"
summary: "Governed Ceph mgr ops: HEALTH_WARN RCA, OSD/PG/pool/RBD/CephFS/RGW, 37 tools."
license: MIT
homepage: https://github.com/AIops-tools/Ceph-AIops
tags: [aiops, mcp, governance, ceph]
description: >
Use this skill whenever the user needs to operate or diagnose a Ceph cluster via its ceph-mgr Dashboard REST API — decode a HEALTH_WARN/ERR state into cause + action (cluster_health), read the cluster status, inspect OSDs (tree/df/perf), placement groups (summary/stuck/scrub), pools (list/usable capacity), RBD images and snapshots, CephFS/MDS and RGW status, monitors/managers, slow ops and capacity forecast — plus governed writes (set cluster flags, reweight/mark-in/mark-out/purge OSDs, trigger scrubs, set pool quota/pg_num/autoscale/size, create/delete pools, create/delete RBD images and snapshots, throttle recovery/backfill).
Always use this skill for "ceph health", "what does this HEALTH_WARN mean", "PG_DEGRADED / OSD_NEARFULL / SLOW_OPS / MON_DOWN", "ceph -s", "which OSD is most full", "drain an OSD", "purge an OSD", "stuck PGs", "overdue scrub", "pool usable capacity", "set pool size / quota", "rebalance is too slow / throttle backfill", "RBD image or snapshot", "MDS behind on trimming", "RGW large omap", "mon quorum", or "days to nearfull" when the context is a Ceph cluster (cephadm, hypervisor-bundled Ceph, or MicroCeph).
Do NOT use when the target is not Ceph — a hypervisor, a different storage appliance, a backup product, a Kubernetes cluster, or a network device. Route those to the appropriate other AIops-tools skill (negative routing hint only).
Common Ceph ops with a built-in governance harness (audit, policy, token budget, undo, risk-tiers).
installer:
kind: uv
package: ceph-aiops
argument-hint: "[ceph question or describe your cluster task]"
allowed-tools:
- Bash
metadata: {"openclaw":{"requires":{"anyBins":["ceph-aiops","uvx"]},"optional":{"env":["CEPH_AIOPS_CONFIG","CEPH_AIOPS_MASTER_PASSWORD"]},"homepage":"https://github.com/AIops-tools/Ceph-AIops","emoji":"🐙","os":["macos","linux"]}}
compatibility: >
Standalone, self-governed Ceph operations. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency. Works against vanilla ceph-mgr (cephadm / hypervisor-bundled Ceph / MicroCeph); no croit and no Kubernetes dependency.
All write operations are audited to a local SQLite DB under ~/.ceph-aiops/ (relocatable via CEPH_AIOPS_HOME).
Connection: the ceph-mgr Dashboard REST API over HTTPS (default port 8443). Authentication is username + password exchanged for a short-lived JWT at POST /api/auth; the mgr 'dashboard' module must be enabled. The username lives in config.yaml; the password is stored ENCRYPTED in ~/.ceph-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'ceph-aiops init' to onboard, or 'ceph-aiops secret set <target>' to_meta.json
{
"ownerId": "kn7b067awq2s97bn3d7p5qfhw5827pxc",
"slug": "ceph-aiops",
"version": "0.11.5",
"publishedAt": 1789601119214
}references/agent-guardrails.md
# Agent guardrails — running ceph-aiops with a smaller / local model
If you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,
Ollama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably
better results with a short system prompt. This page gives you one, and — more
importantly — tells you which guardrails you **no longer need to write**, because
the tool now enforces them itself.
The distinction matters. A guardrail in a prompt is a request. A guardrail in the
harness is a guarantee. Anything below that we could move into the harness, we did.
## Authorization is not this tool's job — decide it where it belongs
Whether a write should happen is your decision, or the account's. The tool does
not gate it — there is no read-only switch and no approval prompt to configure.
The two right places to control read vs write:
- **The account you connect with.** Give it a ceph-mgr Dashboard account with a
read-only role. A write then fails at the mgr, which is the only place the
permission actually lives — no skill-side flag can be argued around by a
model, but a revoked permission cannot be.
- **Your agent's system prompt.** If you want an observe-only session, tell the
model not to call the write tools (they are clearly tagged `[WRITE]`).
What the tool *does* guarantee is that you can always see what happened:
## What the tool enforces — do not waste prompt budget on these
| You might be tempted to prompt | Why you don't need to |
|---|---|
| "Log everything you do, over both MCP and the CLI" | Every call is audited to `~/.ceph-aiops/audit.db` regardless of what the model says it did — and the CLI writes the same row the MCP path does, so there is no unaudited entry point. Reversible writes also record an undo token capturing the *prior* state. |
| "Don't invent a value when a field is missing" | A field the Dashboard did not return comes back as `null`, never as `""`. A missing `deviceClass`, `host`, MDS `state`, or `pg_autoscale_mode` is distinguishable from an empty one in the payload. |
| "Tell me if the output was cut off" | `pg_dump_stuck` returns `{"stuck": [...], "returned": N, "limit": L, "truncated": true/false}` and `pg_summary` the same shape under `unhealthy` (plus `states`) — the list key differs, so look it up per tool rather than expecting `stuck` everywhere. Truncation is measured (one extra row is collected), not guessed. `pg_summary` also keeps `unhealthyCount` as the true total even when the list is capped. |
| "Explain what HEALTH_WARN means" | `cluster_health` already folds each active check code (`PG_DEGRADED`, `OSD_NEARFULL`, `SLOW_OPS`, `LARGE_OMAP_OBJECTS`, …) into a plain-language `cause` and `suggestedAction`. The model should quote those, not compose its own. |
| "Confirm before anything destructive" | Every destructive tool (`osd_purge`, `osd_mark_out`, `pool_delete`, `rbd_image_delete`, `rbd_snapshot_delete`, `set_pool_size`) takes `dry_run=True` for a preview and is `risk=references/capabilities.md
# ceph-aiops capabilities
> 37 MCP tools (17 read, 18 write, 2 undo) over the **ceph-mgr
> Dashboard REST API** (`https://<host>:8443`, JWT via `POST /api/auth`).
> Multi-node rebalance behaviour and the write ops need live verification
> (see `docs/VERIFICATION.md`).
## Read tools (17)
| Tool | API path | Returns |
|------|----------|---------|
| `cluster_health` | `GET /api/health/full` | **flagship RCA** — per active HEALTH_WARN/ERR check: code, plain-language meaning, likely cause, suggested action |
| `cluster_status` | `GET /api/health/minimal` | `ceph -s` summary: health status, mon/mgr/osd/pg counts |
| `osd_tree` | `GET /api/osd` | OSD tree: up/in, CRUSH weight, host, device class |
| `osd_df` | `GET /api/osd` | per-OSD utilization %, **most-full first**, near-full / backfill-full flags |
| `osd_perf` | `GET /api/osd` | commit/apply latency per OSD, **slowest first** |
| `pg_summary` | `GET /api/pg` (+ `/api/health/full`) | PG **state histogram** + list of non-active+clean PGs |
| `pg_dump_stuck` | `GET /api/pg` | stuck PGs (inactive/unclean/stale/undersized) + implicated OSDs |
| `scrub_status` | `GET /api/health/full` | PGs overdue for scrub / deep-scrub |
| `pool_ls` | `GET /api/pool` | pools: name, id, size, pg_num, autoscale mode, application |
| `pool_df` | `GET /api/pool` | per-pool usage; **usable capacity = raw ÷ size** |
| `rbd_ls` | `GET /api/block/image` | RBD images (optionally filtered by pool): name, size, pool |
| `cephfs_status` | `GET /api/cephfs` | MDS ranks + **"behind on trimming"** + client count |
| `rgw_status` | `GET /api/rgw/daemon` + `GET /api/rgw/bucket` | RGW daemons + buckets + **LARGE_OMAP / unsharded-index** findings |
| `mon_status` | `GET /api/monitor` | monitors: in-quorum vs **out-of-quorum** |
| `mgr_status` | `GET /api/health/full` | active mgr, standbys, enabled modules |
| `slow_ops` | `GET /api/health/full` | blocked / slow requests grouped **by OSD** |
| `capacity_forecast` | `GET /api/osd` (+ df) | raw/used/avail + **days-to-nearfull** projection |
## Write tools (18)
| Tool | Risk | API path | Undo / safety |
|------|------|----------|---------------|
| `cluster_flag_set` | medium | `GET`+`PUT /api/osd/flags` | set/unset noout/noscrub/nobackfill/norecover; captures prior flag set (undo) |
| `osd_reweight` | medium | `POST /api/osd/{id}/reweight` | 0.0 = drain; captures prior weight (undo) |
| `osd_mark_in` | medium | `POST /api/osd/{id}/mark` | captures prior up/in state (undo) |
| `osd_mark_out` | **high** | `POST /api/osd/{id}/mark` | drains data; CLI double-confirm + dry-run; captures prior state |
| `osd_purge` | **high** | `DELETE /api/osd/{id}` | destroy + crush rm + auth del; **irreversible**; dry-run + double-confirm |
| `trigger_scrub` | medium | `POST /api/pg/{pgid}/scrub` | schedule a shallow scrub; no prior state |
| `trigger_deep_scrub` | medium | `POST /api/pg/{pgid}/deep_scrub` | schedule a deep (data-integrity) scrub |
| `set_pool_quota` | medium | `PUT /api/pool/{name}` | references/cli-reference.md
# ceph-aiops CLI reference > The CLI is a convenience subset; the full 37-tool surface > is via the MCP server (`ceph-aiops mcp`). Talks to the ceph-mgr Dashboard REST > API (`https://<host>:8443`, JWT via `POST /api/auth`). ## Setup & diagnostics ```bash ceph-aiops init # interactive onboarding wizard ceph-aiops doctor [--skip-auth] # config + secret store + JWT login + mgr-dashboard reachability ceph-aiops mcp # start the MCP server (stdio transport) ``` ## Secrets (encrypted store ~/.ceph-aiops/secrets.enc) ```bash ceph-aiops secret set <target> [--value <password>] # store Dashboard password (hidden prompt if no --value) ceph-aiops secret list # names only — values never shown ceph-aiops secret rm <target> ceph-aiops secret migrate # import legacy plaintext .env (CEPH_<T>_PASSWORD) ceph-aiops secret rotate-password # re-encrypt under a new master password ``` ## Read commands ```bash ceph-aiops overview [--target <t>] # HEALTH status + active checks + OSD up/in ceph-aiops health detail # decode active HEALTH_WARN/ERR checks → cause + action (RCA) ceph-aiops health status # ceph -s summary ceph-aiops osd tree # OSD tree: up/in, weight, host, device class ceph-aiops osd df # per-OSD utilization, most-full first, near/backfill-full flags ``` ## Write commands (governed; risk tier in parentheses) ```bash ceph-aiops osd reweight <osd_id> <weight> [--dry-run] # (med) 0.0 = drain; reversible → prior weight ceph-aiops osd out <osd_id> [--dry-run] # (high) mark out — drains data; double confirm ceph-aiops osd purge <osd_id> [--dry-run] # (high) purge — irreversible; double confirm ``` The remaining writes (cluster flags, pool quota/pg_num/autoscale/size/create/delete, RBD image/snapshot create/delete, trigger scrubs, throttle recovery) are exposed through the **MCP server**, not the CLI. ## Common options - `--target, -t <name>` — target name from `config.yaml` (omit to use the default/first target) - `--dry-run` — print the API call that would be made, change nothing - Destructive commands (`osd out`, `osd purge`) require `--dry-run` review + double confirmation - `doctor --skip-auth` — skip the JWT login / connectivity check (config + secret-store checks only) ## Approver env vars (high-risk ops) ```bash export CEPH_AUDIT_APPROVED_BY='[email protected]' export CEPH_AUDIT_RATIONALE='draining failed OSD 7 per ticket OPS-123' ```
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/zw008/skills/ceph-aiops",
"sourceUrl": "https://clawhub.ai/zw008/skills/ceph-aiops",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T21:32:55.609Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-ceph-aiops/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-ceph-aiops/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T21:32:55.609Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2K downloads",
"href": "https://clawhub.ai/zw008/ceph-aiops",
"sourceUrl": "https://clawhub.ai/zw008/ceph-aiops",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T21:32:55.609Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.11.5",
"href": "https://clawhub.ai/zw008/ceph-aiops",
"sourceUrl": "https://clawhub.ai/zw008/ceph-aiops",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-16T23:25:19.214Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-ceph-aiops/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zw008-ceph-aiops/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.11.5",
"description": "- Removed the obsolete skill-card.md file. - Updated references/agent-guardrails.md documentation. - No changes to functionality or user-facing features.",
"href": "https://clawhub.ai/zw008/ceph-aiops",
"sourceUrl": "https://clawhub.ai/zw008/ceph-aiops",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-16T23:25:19.214Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
