{"id":"17adece5-ccd3-4cb6-b514-7f603e86de04","entityType":"agent","slug":"clawhub-zw008-observability-aiops","name":"observability-aiops","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-observability-aiops","canonicalPath":"/agent/clawhub-zw008-observability-aiops","generatedAt":"2026-10-10T13:32:00.088Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":null},"description":"Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager alerts + silences, Grafana dashboards/datasources/folders, bounded Loki LogQL log reads (labels, query, error-tail), five flagship analyses (firing-alert RCA, target-scrape-health, alert-noise/flap, log-error-burst RCA, log-volume/cardinality) plus an alert->log cross-signal, and guarded writes (create/expire silence, create annotation, update/delete dashboard, reload Prometheus config). Always use this skill for \"Prometheus\", \"PromQL\", \"Alertmanager\", \"Grafana\", \"Loki\", \"LogQL\", \"logs\", \"which targets are down\", \"scrape failing\", \"why is this alert firing\", \"root cause this alert\", \"firing alerts\", \"silence this alert\", \"noisy alerts\", \"alert flapping\", \"recording rule\", \"alerting rule\", \"dashboard\", \"datasource health\", \"reload prometheus config\", \"TSDB cardinality\", \"error burst\", \"log volume\", \"log cardinality\", \"tail errors\" when the context is a self-hosted metrics/logs/observability stack. Do NOT use when the target is something other than a Prometheus/Grafana observability stack (a hypervisor, storage appliance, backup product, container-orchestrator control plane, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope. Governed observability operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (RCAs, governed writes, undo); the Loki surface has not (see docs/VERIFICATION.md).","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.5K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:observability-aiops","sourceUrl":"https://clawhub.ai/zw008/observability-aiops","homepage":"https://clawhub.ai/zw008/skills/observability-aiops","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/observability-aiops","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/observability-aiops","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":63,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"observability-aiops technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":null},"stars":null,"forks":null,"downloads":1462,"packageName":null,"latestVersion":"0.10.4","tractionLabel":"1.5K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T11:32:30.020Z","lastCrawledAt":"2026-10-10T11:32:30.020Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T11:32:30.020Z","lastVerifiedAt":null,"highlights":[{"version":"0.10.4","createdAt":"2026-09-16T23:26:23.574Z","changelog":"## observability-aiops 0.10.4 - Updated documentation in `references/agent-guardrails.md`. - Removed the redundant `skill-card.md` file.","fileCount":7,"zipByteSize":21687},{"version":"0.10.3","createdAt":"2026-09-15T06:12:33.053Z","changelog":"## observability-aiops 0.10.3 changelog - Removed the sample skill-card.md file. - No functional or user-facing changes; this is a documentation/packaging cleanup.","fileCount":7,"zipByteSize":21205},{"version":"0.10.2","createdAt":"2026-09-12T14:37:04.077Z","changelog":"- Documentation update: improved or clarified SKILL.md content. - Removed obsolete or redundant file: skill-card.md. - No functional or feature changes to the skill logic.","fileCount":7,"zipByteSize":21133},{"version":"0.10.1","createdAt":"2026-09-12T10:20:49.283Z","changelog":"- Documentation updates in SKILL.md. - Removed redundant skill-card.md file. - No functional or feature changes; housekeeping only.","fileCount":7,"zipByteSize":21370},{"version":"0.10.0","createdAt":"2026-09-12T01:08:27.922Z","changelog":"- Updated internal metadata requirements and options for bin and env setup (added \"anyBins\" with \"observability-aiops\" or \"uvx\"; made required env vars optional). - Removed reference to the deleted skill-card.md file. - No functional changes to user-facing features or documented capabilities.","fileCount":7,"zipByteSize":21118},{"version":"0.9.0","createdAt":"2026-08-10T06:53:14.818Z","changelog":"- Removed the skill-card.md file. - No changes to functional code or user-facing features. - Documentation footprint reduced.","fileCount":7,"zipByteSize":21174},{"version":"0.8.0","createdAt":"2026-08-03T05:54:14.430Z","changelog":"- Removed the redundant sample file skill-card.md to reduce duplication. - No changes to the core functionality, code, description, or compatibility. - Documentation and user experience remain unchanged.","fileCount":7,"zipByteSize":21078},{"version":"0.7.0","createdAt":"2026-08-02T09:40:58.352Z","changelog":"- Removed the skill-card.md file. - No user-facing features or functionality were changed.","fileCount":7,"zipByteSize":21109}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:observability-aiops","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:observability-aiops` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/observability-aiops before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T13:32:00.084Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-observability-aiops/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":null},"readme":"Skill: observability-aiops\n\nOwner: zw008\n\nSummary: Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager alerts + silences, Grafana dashboards/datasources/folders, bounded Loki LogQL log reads (labels, query, error-tail), five flagship analyses (firing-alert RCA, target-scrape-health, alert-noise/flap, log-error-burst RCA, log-volume/cardinality) plus an alert->log cross-signal, and guarded writes (create/expire silence, create annotation, update/delete dashboard, reload Prometheus config). Always use this skill for \"Prometheus\", \"PromQL\", \"Alertmanager\", \"Grafana\", \"Loki\", \"LogQL\", \"logs\", \"which targets are down\", \"scrape failing\", \"why is this alert firing\", \"root cause this alert\", \"firing alerts\", \"silence this alert\", \"noisy alerts\", \"alert flapping\", \"recording rule\", \"alerting rule\", \"dashboard\", \"datasource health\", \"reload prometheus config\", \"TSDB cardinality\", \"error burst\", \"log volume\", \"log cardinality\", \"tail errors\" when the context is a self-hosted metrics/logs/observability stack. Do NOT use when the target is something other than a Prometheus/Grafana observability stack (a hypervisor, storage appliance, backup product, container-orchestrator control plane, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope. Governed observability operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (RCAs, governed writes, undo); the Loki surface has not (see docs/VERIFICATION.md).\n\nTags: latest:0.10.4\n\nVersion history:\n\nv0.10.4 | 2026-09-16T23:26:23.574Z | auto\n\n## observability-aiops 0.10.4\n\n- Updated documentation in `references/agent-guardrails.md`.\n- Removed the redundant `skill-card.md` file.\n\nv0.10.3 | 2026-09-15T06:12:33.053Z | auto\n\n## observability-aiops 0.10.3 changelog\n\n- Removed the sample skill-card.md file.\n- No functional or user-facing changes; this is a documentation/packaging cleanup.\n\nv0.10.2 | 2026-09-12T14:37:04.077Z | auto\n\n- Documentation update: improved or clarified SKILL.md content.\n- Removed obsolete or redundant file: skill-card.md.\n- No functional or feature changes to the skill logic.\n\nv0.10.1 | 2026-09-12T10:20:49.283Z | auto\n\n- Documentation updates in SKILL.md.\n- Removed redundant skill-card.md file.\n- No functional or feature changes; housekeeping only.\n\nv0.10.0 | 2026-09-12T01:08:27.922Z | auto\n\n- Updated internal metadata requirements and options for bin and env setup (added \"anyBins\" with \"observability-aiops\" or \"uvx\"; made required env vars optional).\n- Removed reference to the deleted skill-card.md file.\n- No functional changes to user-facing features or documented capabilities.\n\nv0.9.0 | 2026-08-10T06:53:14.818Z | auto\n\n- Removed the skill-card.md file.\n- No changes to functional code or user-facing features.\n- Documentation footprint reduced.\n\nv0.8.0 | 2026-08-03T05:54:14.430Z | auto\n\n- Removed the redundant sample file skill-card.md to reduce duplication.\n- No changes to the core functionality, code, description, or compatibility.\n- Documentation and user experience remain unchanged.\n\nv0.7.0 | 2026-08-02T09:40:58.352Z | auto\n\n- Removed the skill-card.md file.\n- No user-facing features or functionality were changed.\n\nv0.6.0 | 2026-07-21T09:42:23.545Z | auto\n\n- Updated and clarified governance mechanism: reduced emphasis on \"policy engine\" and \"risk-tier gate,\" now focusing on budget guard, audit, undo, and risk-tier tagging/labels.\n- Removed or reworded mentions of strict policy and pre-check requirements from documentation.\n- Documentation refreshed for accuracy and to better align with current implementation in SKILL.md and reference files.\n- Removed deprecated skill-card.md file.\n- General clarification and simplification in references to align with latest feature scope.\n\nv0.5.0 | 2026-07-20T11:16:45.255Z | auto\n\n- skill-card.md file removed.  \n- No new features or interface changes.  \n- All core functionality and documentation unchanged.\n\nv0.4.0 | 2026-07-19T03:52:43.048Z | auto\n\n**Expanded capabilities, new documentation, and updated governance for self-hosted observability operations.**\n\n- Added agent guardrails documentation.\n- Increased supported tool count from 37 to 39, with updated tables and overviews.\n- Clarified and enhanced governance, verification, and credential storage descriptions.\n- Updated compatibility and verification notes: Prometheus/Alertmanager/Grafana surfaces exercised live; Loki surface remains mock-only.\n- Removed outdated skill card and performed significant documentation clean-up and refinement.\n\nv0.3.0 | 2026-07-17T05:56:41.485Z | auto\n\n**Version 0.3.0 — Adds Grafana Loki support and log analytics**\n\n- Introduced read-only support for Grafana Loki (logs), including bounded LogQL queries, label exploration, and error-tail analysis.\n- Added five flagship analyses: log error-burst RCA, log volume/cardinality, and alert→log cross-signal, alongside existing Prometheus/Grafana analytics.\n- Expanded “What This Skill Does” matrix and platform compatibility to cover logs as well as metrics.\n- Updated documentation and CLI reference to reflect new Loki and log-analysis capabilities.\n- Removed deprecated skill-card.md file.\n\nv0.2.0 | 2026-07-13T13:10:07.798Z | auto\n\nobservability-aiops 0.2.0\n\n- Documentation updated: SKILL.md refined and expanded for clarity.\n- skill-card.md file removed.\n- No core logic or functionality changes—this release focuses on documentation improvements and cleanup.\n\nv0.1.0 | 2026-07-13T06:24:42.524Z | auto\n\nInitial release – governed operations for self-hosted Prometheus, Alertmanager, and Grafana deployments.\n\n- Provides read/write tools for querying Prometheus (including PromQL), Alertmanager, and Grafana APIs.  \n- Bundles a governance harness: audit log, policy engine, token/runaway budget guard, undo, and risk-tiering.\n- Stores all credentials securely encrypted (never plaintext on disk); supports migration from legacy env-var tokens.\n- Supports both read operations (metrics, targets, rules, alerts, dashboards) and guarded write actions (create/expire silence, update/delete dashboard, reload Prometheus config).\n- Offers analyses: firing-alert root cause, target scrape health, and alert noise/flap identification.\n- Requires no external skill-family; everything is in a standalone package.\n- Preview-only: operations are mock-validated for initial release.\n\nArchive index:\n\nArchive v0.10.4: 7 files, 21687 bytes\n\nFiles: references/agent-guardrails.md (7716b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (3219b), SKILL.md (20269b), _meta.json (139b)\n\nFile v0.10.4:SKILL.md\n\n---\nname: observability-aiops\nslug: observability-aiops\ndisplayName: \"Observability AIops\"\nsummary: \"Governed Prometheus + Grafana ops: PromQL, alerts, dashboards, RCA; 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Observability-AIops\ntags: [aiops, mcp, governance, observability]\ndescription: >\n  Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager alerts + silences, Grafana dashboards/datasources/folders, bounded Loki LogQL log reads (labels, query, error-tail), five flagship analyses (firing-alert RCA, target-scrape-health, alert-noise/flap, log-error-burst RCA, log-volume/cardinality) plus an alert->log cross-signal, and guarded writes (create/expire silence, create annotation, update/delete dashboard, reload Prometheus config).\n  Always use this skill for \"Prometheus\", \"PromQL\", \"Alertmanager\", \"Grafana\", \"Loki\", \"LogQL\", \"logs\", \"which targets are down\", \"scrape failing\", \"why is this alert firing\", \"root cause this alert\", \"firing alerts\", \"silence this alert\", \"noisy alerts\", \"alert flapping\", \"recording rule\", \"alerting rule\", \"dashboard\", \"datasource health\", \"reload prometheus config\", \"TSDB cardinality\", \"error burst\", \"log volume\", \"log cardinality\", \"tail errors\" when the context is a self-hosted metrics/logs/observability stack.\n  Do NOT use when the target is something other than a Prometheus/Grafana observability stack (a hypervisor, storage appliance, backup product, container-orchestrator control plane, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n  Governed observability operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (RCAs, governed writes, undo); the Loki surface has not (see docs/VERIFICATION.md).\ninstaller:\n  kind: uv\n  package: observability-aiops\nargument-hint: \"[a PromQL query, an alert/dashboard uid, or describe your observability task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"observability-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"OBSERVABILITY_AIOPS_CONFIG\",\"OBSERVABILITY_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Observability-AIops\",\"emoji\":\"📈\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed observability operations across Prometheus (HTTP API + PromQL, default port 9090, optional bearer token), a companion Alertmanager (/api/v2, default port 9093), Grafana (HTTP API, default port 3000, required bearer token), and Grafana Loki (HTTP API, default port 3100, optional bearer or basic auth, optional multi-tenant X-Scope-OrgID). Loki is READ-ONLY: bounded LogQL reads only (labels, label values, query_range with a hard lookback + line cap and a stream-selector gate, a canned error-tail), with no write surface. Each target in the config names its own platform, so one config can span the whole stack. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.observability-aiops/ (relocatable via OBSERVABILITY_AIOPS_HOME).\n  Credentials: the Grafana service-account/API token (required) or the Prometheus bearer token (optional; self-hosted Prometheus is often unauthenticated) is stored ENCRYPTED in ~/.observability-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'observability-aiops init' to onboard (it asks for the platform), or 'observability-aiops secret set <target>' to add one. The store is unlocked by a master password from OBSERVABILITY_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var OBSERVABILITY_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback with a deprecation warning (migrate with 'observability-aiops secret migrate'). The token is sent as an Authorization: Bearer header and held only in memory; secrets are never logged or echoed.\n  PromQL is used only through read endpoints (/api/v1/query, /query_range) — there is no write query path. State-changing operations pass through the @governed_tool decorator (budget guard + audit + risk-tier tagging). The destructive write (delete_dashboard) is high-risk with dry_run and captures the full prior dashboard model BEFORE deleting; reversible writes (update_dashboard, create_silence) capture the real fetched before-state and record an inverse undo descriptor. Silences are TIME-BOXED (create_silence requires a positive duration).\n  Webhooks: none — no outbound network calls beyond the configured Prometheus / Alertmanager / Grafana endpoints.\n  SSL: verify_ssl defaults to true; disable for self-signed lab certs.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Verification status: the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (reads, the three metric RCAs, silence + dashboard governed writes, and undo replay); the Loki surface is mock-only so far. Prometheus, Grafana and Loki are free/open-source (docker run prom/prometheus, grafana/grafana, grafana/loki) so a live 'doctor' check is easy. docs/VERIFICATION.md records what was and was not covered.\n---\n\n# Observability AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the Prometheus or Grafana projects, Grafana Labs, or the CNCF.** Prometheus, Alertmanager and Grafana are trademarks of their respective owners. Source at [github.com/AIops-tools/Observability-AIops](https://github.com/AIops-tools/Observability-AIops) under the MIT license.\n\nGoverned self-hosted observability operations — **39 MCP tools** across\n**Prometheus** (HTTP API + PromQL), **Alertmanager** (alerts + silences),\n**Grafana** (dashboards, datasources, folders), and **Grafana Loki** (bounded\nLogQL log reads + log RCA, read-only), every one wrapped with the bundled\n`@governed_tool` harness: a local unified audit log under\n`~/.observability-aiops/`, token/runaway budget guard, undo-token\nrecording, and descriptive risk-tier labels. One config can span the whole\nstack. Bearer tokens are stored **encrypted** (`~/.observability-aiops/secrets.enc`,\nFernet + scrypt) — never plaintext on disk.\n\nThis is the **self-hosted-observability** complement to enterprise monitoring\nsuites: it speaks the open Prometheus/Grafana APIs an SRE actually runs.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`observability_aiops.governance`) — no external skill-family dependency.\n> Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been\n> exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki\n> surface has not yet been exercised live (see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **Metrics** | Prometheus | instant_query, range_query, label_values, series_metadata | 4 | read |\n| **Targets & status** | Prometheus | list_targets, target_scrape_health, dropped_targets, prometheus_config_status, prometheus_tsdb_status | 5 | read |\n| **Rules** | Prometheus | list_rules, rule_health | 2 | read |\n| **Alerts** | Prometheus/Alertmanager | firing_alerts, pending_alerts, alertmanager_alerts, list_silences | 4 | read |\n| **Grafana** | Grafana | list_dashboards, get_dashboard, list_datasources, datasource_health, list_folders | 5 | read |\n| **Loki** | Loki | loki_labels, loki_label_values, loki_query, loki_tail_errors | 4 | read |\n| **Overview + analyses** | all | observability_overview + firing_alert_rca, target_scrape_health_analysis, alert_noise_and_flap_analysis | 4 | read |\n| **Log analyses + cross-signal** | Loki (+ Prometheus) | log_error_burst_rca, log_volume_analysis, alert_log_context | 3 | read |\n| **Writes** | Alertmanager/Grafana/Prometheus | create_silence, expire_silence (med) · create_annotation (med) · update_dashboard (med) · delete_dashboard (**high**) · reload_prometheus_config (med) | 6 | write |\n\nThe three metric flagship analyses are transparent heuristics that report their\nnumbers: `firing_alert_rca` joins each firing alert to its rule expression and\nmaps it to a cause + action; `target_scrape_health_analysis` ranks down/erroring\nscrape targets and classifies each `lastError`; `alert_noise_and_flap_analysis`\nfinds noisy/duplicate alerts and recommends a dedup/rollup. The two **log**\nanalyses mirror this: `log_error_burst_rca` compares per-stream error counts\nagainst a baseline window and classifies each burst (new signature / volume spike\n/ single-instance); `log_volume_analysis` ranks the highest-volume streams and\nwarns on high-cardinality (high-churn) labels. `alert_log_context` bridges the two\nsignals — it maps a firing Prometheus alert's labels to a Loki stream selector and\npulls the correlated logs. **Loki is read-only** (no safe write surface).\n\n## Quick Install\n\n```bash\nuv tool install observability-aiops\nobservability-aiops init       # wizard: pick platform (prometheus/grafana) + encrypted token\nobservability-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/observability-aiops\nopenclaw skills info observability-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a snapshot (`overview` / `observability_overview`): firing-alert count,\n  scrape targets up/down, rules erroring (Prometheus) or dashboard/datasource\n  counts (Grafana)\n- Run PromQL (`instant_query` / `range_query`), enumerate `label_values` or\n  `series_metadata`\n- Check scrape health (`target_scrape_health`, `dropped_targets`) and rule health\n  (`rule_health`, `list_rules`)\n- Triage alerts: `firing_alerts` / `pending_alerts`, the Alertmanager view\n  (`alertmanager_alerts`, `list_silences`), then `firing_alert_rca` to root-cause\n- Reduce alert noise (`alert_noise_and_flap_analysis`) → group_by / inhibition /\n  longer `for`\n- Grafana: `list_dashboards`, `get_dashboard`, `list_datasources`,\n  `datasource_health`, `list_folders`\n- Loki logs: enumerate `loki_labels` / `loki_label_values`, run a bounded\n  `loki_query` (LogQL, stream selector required), `loki_tail_errors` for a\n  selector; then `log_error_burst_rca` to root-cause an error burst and\n  `log_volume_analysis` for volume/cardinality; `alert_log_context` to pull the\n  logs behind a firing alert\n- Governed writes: silence an alert (`create_silence`, time-boxed), annotate an\n  event (`create_annotation`), update/delete a dashboard (`dry_run` first for\n  either), or hot-reload Prometheus (`reload_prometheus_config`)\n\n**Do NOT use when** the target is not a Prometheus/Grafana observability stack —\nroute hypervisor, storage, backup, container-orchestrator, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill. Hosted/SaaS\nmonitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Prometheus / Alertmanager / Grafana observability ops | **observability-aiops** (this skill) |\n| A different platform (hypervisor, storage, backup, orchestrator, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Hosted/SaaS monitoring (Datadog, New Relic, enterprise NMS) | out of scope for this tool |\n\n## Common Workflows\n\n> The CLI covers the reads and the three RCAs (`alert`, `query`, `logs`,\n> `overview`); the guarded **writes** (silences, annotations, dashboards,\n> config reload) are MCP tools — those steps name the tool rather than a CLI\n> command.\n\n### \"Pager went off\" — root-cause the firing alerts and time-box the noise\n\n1. `observability-aiops overview` → one-shot stack picture: firing counts, target\n   health, rule health — is this one alert or the whole stack?\n2. `observability-aiops alert firing` → what is firing right now, grouped by\n   severity\n3. `observability-aiops alert rca` → each firing alert joined to its rule\n   expression with a likely cause and a recommended action (advisory heuristic —\n   verify it, do not act on it blind)\n4. `observability-aiops query instant '<the rule expr>'` → evaluate the alert's\n   own expression yourself and confirm the RCA's reading of it\n5. `observability-aiops query range '<expr>' --start <rfc3339> --end <rfc3339> --step 60s`\n   → see when it crossed the threshold, which usually names the change that\n   caused it\n6. Time-box the noise while you fix the cause: the `create_silence` MCP tool on a\n   specific matcher (a positive duration is **required** — silences cannot be\n   open-ended), then `observability-aiops alert silences` to confirm it landed\n7. **Failure branch**: if the silence was too broad, `expire_silence` ends it\n   immediately, or `observability-aiops undo apply <id>` replays the recorded\n   inverse (`create_silence`'s undo is expire). If `alert rca` returns nothing\n   while alerts are visibly firing, the alerts are coming from Alertmanager\n   without a matching Prometheus rule — check `alertmanager_alerts` and\n   `list_rules` rather than assuming the RCA is broken.\n\n### Investigate a scrape gap (\"metrics went missing\")\n\n1. `observability-aiops overview` → up/down target counts at a glance\n2. `target_scrape_health` → the unhealthy targets with their raw `lastError`\n3. `target_scrape_health_analysis` → down targets ranked, each `lastError`\n   classified (connection refused / timeout / auth / DNS / TLS) with a concrete fix\n4. `dropped_targets` → if a target is missing **entirely** rather than down, it\n   was relabeled away; this is where that shows up\n5. `observability-aiops query instant 'up{job=\"<job>\"}'` → confirm the gap in the\n   metric itself, not just in the target page\n6. After fixing scrape config, `reload_prometheus_config` (a governed write) →\n   then re-run `target_scrape_health` to confirm the target came back\n7. **Failure branch**: if `reload_prometheus_config` succeeds but the target is\n   still down, the config on disk was not what you thought — check\n   `prometheus_config_status` for what Prometheus actually loaded. A reload with\n   a broken config is rejected by Prometheus and leaves the old config running,\n   so a failed reload is not an outage.\n\n### Tame a noisy / flapping alert\n\n1. `observability-aiops alert firing` → the volume of what is firing\n2. `alert_noise_and_flap_analysis` → alertnames with many instances or exact\n   duplicates, each with a `group_by` / inhibition / longer-`for` recommendation\n3. `list_rules` and `rule_health` → read the offending rule's current `for`\n   duration and confirm it is evaluating cleanly\n4. `observability-aiops query range '<rule expr>' --start <rfc3339> --end <rfc3339> --step 60s`\n   → see the flapping in the data and pick a `for` window that actually covers it\n5. `create_silence` for a time-boxed quiet period while the rule change ships;\n   `observability-aiops alert silences` to confirm\n6. **Failure branch**: silencing is a stopgap, not a fix — if the silence expires\n   and the flapping returns, the rule threshold or `for` window is still wrong.\n   Use `observability-aiops undo list` to see exactly which silences this tool\n   created, so no stale silence quietly hides a real outage.\n\n### Root-cause a log error burst (Loki, read-only)\n\n1. `alert_log_context <alertname>` → the firing alert's labels mapped to a Loki\n   stream selector plus the correlated error logs (or start from a selector directly)\n2. `observability-aiops logs errors '{app=\"api\"}' --hours 2 --limit 200` → tail\n   the error-level lines for that stream\n3. `log_error_burst_rca <selector>` → per-stream error counts against a baseline\n   window, each burst classified (new signature / volume spike / single instance)\n4. `observability-aiops logs query '{app=\"api\"} |= \"timeout\"' --hours 2` → confirm\n   the specific signature the RCA named\n5. `log_volume_analysis <selector>` → the highest-volume streams and any\n   high-cardinality label driving a stream/index explosion\n6. **Failure branch**: Loki here is **read-only and bounded** — queries require a\n   stream selector and are capped by lookback and line count. A query rejected\n   for a missing selector is the guard working, not a bug: narrow it with\n   `observability-aiops logs labels` first. There is no write surface for Loki,\n   so remediation happens in the emitting service, not through this tool.\n\n### Safely change or retire a Grafana dashboard (reversible)\n\n1. `list_dashboards` / `list_folders` → locate the dashboard and its folder\n2. `get_dashboard <uid>` → confirm this is the right dashboard before touching it\n3. `update_dashboard` with `dry_run=True` → preview; then for real — it fetches\n   and stashes the **prior model** and records a restore undo\n4. To retire one: `delete_dashboard <uid>` with `dry_run=True` first. Delete is\n   `high` risk — the prior model is captured **before** the delete so the undo\n   can recreate it; set `OBSERVABILITY_AUDIT_APPROVED_BY` (and\n   `OBSERVABILITY_AUDIT_RATIONALE`) if you want that recorded on the audit row\n5. `create_annotation` → mark the change on the timeline so the next responder\n   can correlate a metric shift with this edit\n6. **Failure branch**: wrong dashboard or a bad edit — `observability-aiops undo list`\n   then `observability-aiops undo apply <id>` restores the captured prior model\n   (or recreates a deleted dashboard from it). If the write fails outright, that\n   is the connecting account's permissions (this tool does not gate it) — check\n   the token's role before assuming `observability-aiops doctor` connectivity is\n   at fault.\n\n## Governance & Safety\n\nThe skill delivers reads and writes and records them; it does **not** decide whether a write is\npermitted. That is your agent's judgement, or the permission of the account you connect it with\n(give it a Grafana token with only Viewer scope, and a Prometheus/Alertmanager reached without the\nadmin/write API — writes then fail at the server). There is no read-only switch, policy file, or\napproval gate.\n\n- **Audit is the guarantee, and it is not bypassable.** Every operation — MCP and CLI alike — is\n  logged to `~/.observability-aiops/audit.db` (relocatable via `OBSERVABILITY_AIOPS_HOME`): params,\n  result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.\n- `OBSERVABILITY_AUDIT_APPROVED_BY` / `OBSERVABILITY_AUDIT_RATIONALE` are optional annotations\n  recorded on the audit row (who/why); they are never required and never block.\n- **Runaway guard** — a safety backstop, not authorization: the same call looped in a tight window\n  trips a circuit breaker. Disable with `OBSERVABILITY_RUNAWAY_MAX=0`.\n- Writes support `--dry-run` / `dry_run=True` and double confirmation at the CLI.\n- Silences are **time-boxed** (require a positive duration). Reversible writes\n  capture the real fetched before-state and record an inverse descriptor\n  (create_silence→expire, update/delete dashboard→restore/recreate).\n\n## References\n\n- `references/capabilities.md` — full tool + platform + API-path reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, credentials, and connectivity\n- `references/agent-guardrails.md` — running this with a smaller / local model:\n  what the harness enforces for you, and a ready-made system prompt for the rest\n\nFile v0.10.4:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"observability-aiops\",\n  \"version\": \"0.10.4\",\n  \"publishedAt\": 1789601183574\n}\n\nFile v0.10.4:references/agent-guardrails.md\n\n# Agent guardrails — running observability-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the API did not return comes back as `null`, never as `\"\"`. An alert with no `severity` label, a scrape target that has never errored, a recording rule with no alert `state`, a silence with no `comment` — all report `null`, distinguishable from a genuinely empty value. |\n| \"Tell me if the output was cut off\" | Bounded reads (`loki_query`, `loki_tail_errors`, `loki_labels`, `loki_label_values`, `instant_query`, `range_query`, `label_values`, `series_metadata`, `undo_list`) return `{\"returned\": N, \"limit\": L, \"truncated\": true/false}` alongside the rows. For the Loki reads truncation is **measured** — one line beyond the limit is requested — not guessed from a length coincidence. |\n| \"Preserve the ordering / tell me what's most urgent\" | The analysis tools already return worst-first: `firing_alert_rca` ranks by severity, `target_scrape_health_analysis` puts down targets before slow ones, `alert_noise_and_flap_analysis` sorts by instance count. Priority is the list order, and each entry carries the measured number it was ranked on. |\n| \"Confirm before anything destructive\" | Every write tool takes `dry_run` for a preview, and `delete_dashboard` is `risk=high`. ⚠️ **Apart from `undo apply`, no write tool here has a CLI command**, so the CLI double confirmation never applies to one: they are reachable only over MCP, where nothing prompts. Keep your own confirmation for them. |\n| \"Log what you did\" | Every call is audited to `~/.observability-aiops/audit.db` regardless of what the model says it did, and reversible writes record an undo token (`undo_list` / `undo_apply`). |\n| \"Don't hammer the same call in a loop\" | The runaway guard trips a circuit breaker on tight poll/retry loops — a safety backstop, not authorization. |\n\nAuthorization is not this tool's job. Whether a write is allowed to happen is\ndecided by the account you connect it with, or by your agent's own judgement —\nnot by this harness. See \"Recommended setup for a local model\" below for how to\nenforce read-only at the account instead of in a prompt.\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a self-hosted observability stack (Prometheus, Alertmanager,\nGrafana, Loki) through the observability-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current state of the stack, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit (or a narrower\n  selector) instead of treating the partial result as complete.\n- A null field means the API did not return that value. Report it as \"not\n  available\" — never infer it. A missing \"severity\" label is not \"info\".\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  alert names, severities, label values, or health strings.\n- When an analysis tool returns ranked findings, work in the order given and\n  cite the measured number the ranking is based on.\n\nQUERIES\n- PromQL goes to instant_query / range_query; LogQL goes to loki_query. They are\n  different languages — do not send one to the other.\n- A LogQL query MUST carry a stream selector (e.g. '{app=\"api\"}'). A query\n  without one is rejected by the tool, not silently widened.\n- Use label_values / loki_labels / loki_label_values to discover real label\n  names and values before writing a selector. Do not guess a job, instance, or\n  app name.\n\n- Apart from `undo apply`, no write tool here has a CLI command, so nothing will ask you\n  to confirm one —\n  `delete_dashboard` included. Call with `dry_run=True` first, show the operator what would\n  change, and wait for an explicit go-ahead.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, saturation, or regression unless a tool result\n  supports it.\n- Do not add generic advice that does not follow from the tool output.\n- Keep the identifiers straight: an alertname is not a label value; a silence ID\n  is not a dashboard UID; a \"job\" is a scrape config name while an \"instance\" is\n  a single scraped endpoint.\n```\n\n## Recommended setup for a local model\n\nThere is no read-only switch to set — this tool does not decide whether a\nwrite is permitted. If you want the connection to be read-only until you trust\nthe setup, enforce it at the account: give it a Grafana token with only Viewer\nscope, and a Prometheus/Alertmanager reached without the admin/write API. Any\nwrite attempt then fails at the server, which is the place that actually owns\nthe permission.\n\n```bash\nobservability-aiops doctor\n```\n\nWhen you are ready to allow writes (silences, annotations, dashboards), connect\nwith a token that has write scope, and optionally name yourself on the audit\nrow — it is an annotation, not a gate:\n\n```bash\nexport OBSERVABILITY_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport OBSERVABILITY_AUDIT_RATIONALE=\"silencing NodeDiskFilling during the 2026-07-20 disk swap\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the analysis tools —\n  `firing_alert_rca`, `target_scrape_health_analysis`,\n  `alert_noise_and_flap_analysis`, `log_error_burst_rca`, and\n  `alert_log_context` do the multi-step correlation inside one call, so the\n  model does not have to chain reads and keep alertnames, jobs, and selectors\n  straight across turns.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions, scope PromQL and LogQL with real label matchers, and use `--limit`\n  deliberately rather than pulling whole label sets or metric-name lists.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\n## Verification status\n\nThese tools have been exercised against a real stack — Prometheus 3.x,\nAlertmanager, and Grafana 13 — covering firing-alert and scrape-target RCA,\ngoverned silence and dashboard writes, and undo replay. The behaviours described\nabove are observed, not only unit-tested.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Observability-AIops](https://github.com/AIops-tools/Observability-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.4:references/capabilities.md\n\n# observability-aiops capability matrix\n\n> **39 MCP tools** (32 read, 7 write) across Prometheus\n> (HTTP API + PromQL, default port 9090, optional bearer token), a companion\n> Alertmanager (`/api/v2`, port 9093), Grafana (HTTP API, port 3000, required\n> bearer token), and Grafana Loki (HTTP API, port 3100, optional bearer/basic\n> auth, optional multi-tenant `X-Scope-OrgID`). Loki is **read-only**. The\n> Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md).\n\n## Metrics — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `instant_query` | `/api/v1/query` | PromQL evaluated at one instant (samples: metric + value + timestamp) |\n| `range_query` | `/api/v1/query_range` | PromQL over a time range (per-series point arrays) |\n| `label_values` | `/api/v1/label/<name>/values` | distinct values of a label (default `__name__` = all metric names) |\n| `series_metadata` | `/api/v1/series` | series (label-set) metadata for a selector |\n\n## Targets & status — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_targets` | `/api/v1/targets` | active scrape targets (job, instance, health, lastError), optional up/down filter |\n| `target_scrape_health` | `/api/v1/targets` | up/down summary + the unhealthy targets |\n| `dropped_targets` | `/api/v1/targets` | targets discovered but dropped by relabeling |\n| `prometheus_config_status` | `/api/v1/status/config` | running-config fingerprint (sha256) + size — never the raw YAML/secrets |\n| `prometheus_tsdb_status` | `/api/v1/status/tsdb` | TSDB head cardinality + top metrics by series count |\n\n## Rules — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_rules` | `/api/v1/rules` | recording + alerting rules (name, type, expr, health), optional type filter |\n| `rule_health` | `/api/v1/rules` | rule-evaluation health summary + erroring rules |\n\n## Alerts — Prometheus + Alertmanager (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `firing_alerts` | `/api/v1/alerts` | firing Prometheus rule alerts, grouped by severity |\n| `pending_alerts` | `/api/v1/alerts` | pending (not-yet-firing) rule alerts |\n| `alertmanager_alerts` | AM `/api/v2/alerts` | alerts as Alertmanager sees them (post grouping/silence/inhibit) |\n| `list_silences` | AM `/api/v2/silences` | silences (active, pending, expired) with matchers |\n\n## Grafana (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_dashboards` | `/api/search?type=dash-db` | dashboards (uid, title, folder, tags), optional title query |\n| `get_dashboard` | `/api/dashboards/uid/{uid}` | one dashboard's summary (title, version, panel + tag counts) |\n| `list_datasources` | `/api/datasources` | datasources (id, uid, name, type, default flag) |\n| `datasource_health` | `/api/datasources/{id}/health` | one datasource's health (status, message) |\n| `list_folders` | `/api/folders` | Grafana folders |\n\n## Loki — logs (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `loki_labels` | `/loki/api/v1/labels` | distinct label names in the lookback window |\n| `loki_label_values` | `/loki/api/v1/label/<name>/values` | distinct values of one label (name percent-encoded) |\n| `loki_query` | `/loki/api/v1/query_range` | bounded LogQL passthrough — **requires a stream selector**; lookback capped at `MAX_LOOKBACK_HOURS=24`, lines clamped to `MAX_LINE_LIMIT=1000` (default 100) |\n| `loki_tail_errors` | `/loki/api/v1/query_range` | canned error-level read for a selector (line-filter on `(?i)(error\\|fatal\\|panic\\|exception\\|traceback\\|stacktrace)`) |\n\nBounding gate: a LogQL query with **no `{…}` stream selector**, an empty query,\nor a lookback beyond the cap is rejected up front with a teaching error — no\nunbounded scan is ever issued. Label values interpolated into a selector are\nbackslash-escaped; label names in a path are percent-encoded. Loki auth is\noptional (bearer or, per target, `basic` with a `user:password` secret) and a\nmulti-tenant `X-Scope-OrgID` header is sent when the target sets `org_id`.\n\n## Overview & flagship analyses (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `observability_overview` | platform-aware | Prometheus: firing count + targets up/down + rules erroring; Grafana: dashboard/datasource/folder counts; Loki: label-name count |\n| `firing_alert_rca` | firing alerts + alerting rules | each firing alert joined to its rule expr, ranked by severity, mapped to a likely **cause + action** |\n| `target_scrape_health_analysis` | active targets | down/erroring scrapes ranked, each `lastError` classified (refused/timeout/auth/DNS/TLS) with a fix |\n| `alert_noise_and_flap_analysis` | alert instances | alertnames with many instances / exact duplicates flagged with a group_by / inhibition / longer-`for` recommendation |\n\n## Loki — log analyses & cross-signal (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `log_error_burst_rca` | selector + window (pulls current + baseline error streams) | per-stream error counts vs a baseline window; each burst classified **new_signature** (baseline 0), **volume_spike** (>= `burst_ratio`×baseline), or **single_instance** (localized to one pod/instance) with a cause + action + sample lines |\n| `log_volume_analysis` | selector + window (pulls streams + `index/stats`) | top streams by line volume, high-cardinality (high-churn) label warnings, and a retention hint from total ingest bytes |\n| `alert_log_context` | firing alertname (Prometheus) + Loki target | maps the alert's labels → a Loki stream selector (intersect with namespace/job/service/app/container/pod/instance/component, first 4 in priority order; values escaped) and returns the correlated error streams. **Best-effort**: only labels the alert and Loki share will match |\n\n## Undo (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `undo_list` | local undo store (`limit`) | recorded, not-yet-applied reversible writes: undoId, ts, originalTool, inverseTool, note |\n\n## Writes (governed)\n\n| Tool | Risk | API path | Notes |\n|------|------|----------|-------|\n| `create_silence` | **med** | AM `POST /api/v2/silences` | **time-boxed** (requires minutes > 0); returns silenceId; undo → `expire_silence` |\n| `expire_silence` | **med** | AM `DELETE /api/v2/silence/{id}` | inverse of create_silence |\n| `create_annotation` | **medium** | `POST /api/annotations` | Grafana event marker |\n| `update_dashboard` | **med** | `POST /api/dashboards/db` | GETs the prior model first → captures it for a restore undo |\n| `delete_dashboard` | **HIGH** | `DELETE /api/dashboards/uid/{uid}` | `dry_run`; captures prior model **BEFORE** delete; undo → recreate |\n| `reload_prometheus_config` | **med** | `POST /-/reload` | records the pre-reload config hash; no undo (re-apply the prior config file) |\n| `undo_apply` | **med** | dispatches the recorded inverse tool | executes a recorded inverse; the inverse runs through its own governed tool (its real risk tier applies); single-use token; supports `dry_run` |\n\n**No Loki writes.** Loki exposes no safe operational write surface here (no\nsilence/annotation analogue), so this tool ships Loki as read-only by design.\n\n## Out of scope (by design)\n\n- **Hosted/SaaS monitoring** — Datadog, New Relic, and enterprise NMS (only\n  self-hosted Prometheus + Grafana + Loki here)\n- **Loki writes / ingestion / deletes** — read-only LogQL only (no push, no\n  delete-series, no ruler/config writes)\n- **Creating/editing Prometheus rules or scrape config**, and provisioning\n  Grafana datasources/dashboards from scratch (beyond update/delete of an existing\n  dashboard)\n- **Long-term-storage query fan-out** (Thanos/Cortex/Mimir) — the single\n  Prometheus HTTP API only\n\nFile v0.10.4:references/cli-reference.md\n\n# observability-aiops CLI reference\n\n> Covers Prometheus (HTTP API + PromQL), a companion Alertmanager, Grafana\n> (HTTP API), and Grafana Loki (LogQL, read-only). The Prometheus/Alertmanager/\n> Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack;\n> the Loki surface has not (see docs/VERIFICATION.md). The CLI is a convenience\n> subset — the full 39-tool surface is via the MCP server\n> (`observability-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nobservability-aiops init                      # interactive wizard (asks for the platform: prometheus/grafana/loki)\nobservability-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   Prometheus: /api/v1/status/buildinfo · Grafana: /api/health\n                                           #   Loki: /ready + /loki/api/v1/status/buildinfo\nobservability-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.observability-aiops/secrets.enc)\n\n```bash\nobservability-aiops secret set <target> [--value <token>]   # store bearer token (hidden prompt if no --value)\nobservability-aiops secret list                             # names only — secrets never shown\nobservability-aiops secret rm <target>\nobservability-aiops secret migrate                          # import legacy plaintext env (OBSERVABILITY_<TARGET>_TOKEN)\nobservability-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nobservability-aiops overview [--target <t>]   # snapshot: firing alerts + targets up/down + rules erroring (Prometheus)\n                                           #   or dashboard/datasource/folder counts (Grafana) / label-name count (Loki)\n```\n\n## Query (Prometheus PromQL)\n\n```bash\nobservability-aiops query instant 'up'                      # PromQL instant query\nobservability-aiops query range 'rate(x[5m])' --start ... --end ... [--step 60s]\nobservability-aiops query labels [__name__]                 # distinct label values (default = all metric names)\n```\n\n## Logs (Grafana Loki, read-only, bounded)\n\n```bash\nobservability-aiops logs labels [--hours 1] [--target <t>]              # distinct Loki label names in the window\nobservability-aiops logs query '{app=\"api\"} |= \"error\"' [--hours 1] [--limit 100]   # bounded LogQL (stream selector required)\nobservability-aiops logs errors '{app=\"api\"}' [--hours 1] [--limit 100]  # canned error-level tail for a selector\n```\n\n## Alerts\n\n```bash\nobservability-aiops alert firing [--target <t>]             # firing Prometheus rule alerts, by severity\nobservability-aiops alert silences [--target <t>]           # Alertmanager silences\nobservability-aiops alert rca [--target <t>]                # root-cause firing alerts (join to rule expr → cause+action)\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `query`, `logs`, and `alert` are the CLI subset; the remaining\n  metrics, targets, rules, Grafana, Loki analyses (log_error_burst_rca,\n  log_volume_analysis, alert_log_context), and governed-write tools\n  (create/expire silence, create annotation, update/delete dashboard, reload\n  config) are exposed through the MCP server. High-risk MCP writes honour `OBSERVABILITY_AUDIT_APPROVED_BY` /\n  `OBSERVABILITY_AUDIT_RATIONALE` and support dry-run.\n\nFile v0.10.4:references/setup-guide.md\n\n# observability-aiops setup & security guide\n\n> The Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md). **Prometheus,\n> Grafana, and Loki are all free/open-source and trivial to stand up in a lab\n> (`docker run prom/prometheus`, `grafana/grafana`, `grafana/loki`), so a live\n> `doctor` check is easy.**\n\n## 1. Install\n\n```bash\nuv tool install observability-aiops\n```\n\n## 2. Get a credential\n\n- **Prometheus** — a bearer token is **optional**; many self-hosted deployments\n  are unauthenticated. observability-aiops talks to the HTTP API on port **9090**\n  and a companion Alertmanager on **9093** (`/api/v2`).\n- **Grafana** — a **service-account token** (Administration → Service accounts →\n  Add token) or legacy API key is **required**. Grafana's HTTP API is on port\n  **3000**.\n- **Loki** — auth is **optional** (bearer token, or per-target `basic` auth where\n  the stored secret is `user:password`). The HTTP API is on port **3100**. For a\n  multi-tenant deployment set `org_id` to send the `X-Scope-OrgID` header. Loki is\n  **read-only** here.\n\n## 3. Onboard\n\n```bash\nobservability-aiops init\n```\n\nThe wizard asks, per target, for the **platform** (`prometheus` / `grafana` /\n`loki`), the **host**, the **scheme** (`http` / `https`), the **port** (defaults\n9090 for Prometheus, 3000 for Grafana, 3100 for Loki), an optional **Alertmanager\nURL** (Prometheus only), the **auth type** and **org id** (Loki only), and the\n**token** — required for Grafana, optional for Prometheus/Loki. Non-secret\nconnection details go to `~/.observability-aiops/config.yaml`; the token is stored\n**encrypted** into `~/.observability-aiops/secrets.enc`. Example config (one config\ncan span the whole stack):\n\n```yaml\ntargets:\n  - name: prod-prom\n    platform: prometheus\n    host: 10.0.0.20\n    scheme: http\n    port: 9090\n    alertmanager_url: http://10.0.0.20:9093   # optional; blank assumes host:9093\n  - name: prod-grafana\n    platform: grafana\n    host: 10.0.0.30\n    scheme: https\n    port: 3000\n    verify_ssl: true\n  - name: prod-loki\n    platform: loki\n    host: 10.0.0.40\n    scheme: http\n    port: 3100\n    auth_type: bearer        # or 'basic' (secret is user:password)\n    org_id: team-a           # optional; sent as X-Scope-OrgID (multi-tenant)\n```\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nExport the master password so the encrypted store can be unlocked without a\nprompt:\n\n```bash\nexport OBSERVABILITY_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## Credential security\n\n- The token is **never** written to disk in plaintext. It lives only in\n  `~/.observability-aiops/secrets.enc`, encrypted with Fernet (AES-128-CBC +\n  HMAC), the key derived from your master password via scrypt. Only a per-store\n  random salt and the ciphertext are on disk (chmod 600); the master password\n  itself is never stored.\n- A legacy plaintext env var `OBSERVABILITY_<TARGET_NAME_UPPER>_TOKEN` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `observability-aiops secret migrate` (it imports then renames the old `.env`).\n- The token is sent as an `Authorization: Bearer` header at request time and held\n  only in memory; it is never logged or echoed. Exception text and tracebacks are\n  scrubbed of secret-shaped strings before being written to the audit log.\n\n## Governance harness state\n\nState lives under `~/.observability-aiops/` (relocate with `OBSERVABILITY_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with risk tier and any operator-supplied\n  approver/rationale (optional annotations, never required)\n- `undo.db` — inverse descriptors for reversible writes (create_silence→expire,\n  update/delete dashboard→restore/recreate)\n- budget / runaway guard — caps cumulative tool calls and wall-time; trips on\n  tight poll/retry loops\n\n## Governed writes\n\n- **High-risk** op (`delete_dashboard`) supports `dry_run` and captures the full\n  prior dashboard model **before** deleting so the recorded undo can recreate\n  it. Optionally set `OBSERVABILITY_AUDIT_APPROVED_BY` and\n  `OBSERVABILITY_AUDIT_RATIONALE` to annotate the audit row — neither is\n  required, and the write runs either way.\n- **Reversible** writes capture the real fetched before-state:\n  `update_dashboard` (restore prior model), `create_silence` (expire the created\n  silence). `reload_prometheus_config` records the pre-reload config hash.\n- **Time-boxed** ops require a positive duration: `create_silence` (in minutes).\n  This prevents forgotten, indefinite silences.\n\n## Verify\n\n```bash\nobservability-aiops doctor\n```\n\n`doctor` is platform-aware: it checks the config file, the encrypted store and its\npermissions, that a token is present where required, and (unless `--skip-auth`)\nconnectivity — `/api/v1/status/buildinfo` for Prometheus targets, `/api/health`\nfor Grafana targets, and `/ready` + `/loki/api/v1/status/buildinfo` for Loki\ntargets.\n\n## Loki query bounding (safety)\n\nLoki reads are deliberately bounded so an agent can't ask for \"all logs, forever\":\n\n- Every `loki_query` must carry a `{…}` **stream selector** — an unbounded query\n  with no selector (or an empty query) is rejected with a teaching error.\n- Lookback is capped at **24h** (`MAX_LOOKBACK_HOURS`); a longer window is refused.\n- Returned lines are clamped to **1000** (`MAX_LINE_LIMIT`, default 100).\n- `loki_tail_errors` wraps a selector with a canned case-insensitive error filter.\n- Label values interpolated into a selector are backslash-escaped and label names\n  in a path segment are percent-encoded, so a hostile label value can't break out\n  of the LogQL string or rewrite the request path.\n\nFile v0.10.4:skill-card.md\n\n## Description:\n\nOperates self-hosted Prometheus, Alertmanager, Grafana, and Grafana Loki environments for observability snapshots, PromQL and LogQL reads, alert and scrape-health analysis, dashboard and datasource inspection, root-cause workflows, and governed operational writes.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nSREs, developers, and operations engineers use this skill to inspect and triage self-hosted Prometheus, Alertmanager, Grafana, and Loki stacks. It supports alert RCA, scrape-health investigation, dashboard and datasource checks, bounded log analysis, and carefully governed changes such as time-boxed silences, annotations, dashboard updates or deletes, and Prometheus reloads.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can perform write-capable operations against Alertmanager, Grafana, and Prometheus, including dashboard deletion and configuration reloads, without an enforced approval gate.\n\nMitigation: Use least-privilege read-only credentials by default, require operator approval before write-capable MCP calls, and run supported dry-runs before changes.\n\nRisk: Write-capable credentials could allow accidental or unauthorized monitoring changes.\n\nMitigation: Avoid dashboard delete or admin/write API permissions unless they are required for the workflow, and scope service accounts to the smallest practical set of capabilities.\n\nRisk: The master password unlocks the local encrypted secret store for non-interactive or MCP use.\n\nMitigation: Protect OBSERVABILITY_AIOPS_MASTER_PASSWORD as a production secret and avoid exposing it to untrusted agents, logs, or shell history.\n\nRisk: Loki support has not been exercised against a live stack according to the artifact documentation.\n\nMitigation: Validate Loki connectivity and expected query behavior with a live doctor check before relying on Loki results in production workflows.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/observability-aiops)\n- [Publisher profile](https://clawhub.ai/user/zw008)\n- [Project homepage](https://github.com/AIops-tools/Observability-AIops)\n- [Capability matrix](references/capabilities.md)\n- [CLI reference](references/cli-reference.md)\n- [Setup and security guide](references/setup-guide.md)\n- [Agent guardrails](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown and structured tool-result summaries with inline shell commands or configuration snippets when relevant]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include bounded observability query results, RCA summaries, audit-oriented write guidance, and dry-run recommendations.]\n\n## Skill Version(s):\n\n0.10.4 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.3: 7 files, 21205 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2647b), SKILL.md (20269b), _meta.json (139b)\n\nFile v0.10.3:SKILL.md\n\n---\nname: observability-aiops\nslug: observability-aiops\ndisplayName: \"Observability AIops\"\nsummary: \"Governed Prometheus + Grafana ops: PromQL, alerts, dashboards, RCA; 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Observability-AIops\ntags: [aiops, mcp, governance, observability]\ndescription: >\n  Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager alerts + silences, Grafana dashboards/datasources/folders, bounded Loki LogQL log reads (labels, query, error-tail), five flagship analyses (firing-alert RCA, target-scrape-health, alert-noise/flap, log-error-burst RCA, log-volume/cardinality) plus an alert->log cross-signal, and guarded writes (create/expire silence, create annotation, update/delete dashboard, reload Prometheus config).\n  Always use this skill for \"Prometheus\", \"PromQL\", \"Alertmanager\", \"Grafana\", \"Loki\", \"LogQL\", \"logs\", \"which targets are down\", \"scrape failing\", \"why is this alert firing\", \"root cause this alert\", \"firing alerts\", \"silence this alert\", \"noisy alerts\", \"alert flapping\", \"recording rule\", \"alerting rule\", \"dashboard\", \"datasource health\", \"reload prometheus config\", \"TSDB cardinality\", \"error burst\", \"log volume\", \"log cardinality\", \"tail errors\" when the context is a self-hosted metrics/logs/observability stack.\n  Do NOT use when the target is something other than a Prometheus/Grafana observability stack (a hypervisor, storage appliance, backup product, container-orchestrator control plane, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n  Governed observability operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (RCAs, governed writes, undo); the Loki surface has not (see docs/VERIFICATION.md).\ninstaller:\n  kind: uv\n  package: observability-aiops\nargument-hint: \"[a PromQL query, an alert/dashboard uid, or describe your observability task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"observability-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"OBSERVABILITY_AIOPS_CONFIG\",\"OBSERVABILITY_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Observability-AIops\",\"emoji\":\"📈\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed observability operations across Prometheus (HTTP API + PromQL, default port 9090, optional bearer token), a companion Alertmanager (/api/v2, default port 9093), Grafana (HTTP API, default port 3000, required bearer token), and Grafana Loki (HTTP API, default port 3100, optional bearer or basic auth, optional multi-tenant X-Scope-OrgID). Loki is READ-ONLY: bounded LogQL reads only (labels, label values, query_range with a hard lookback + line cap and a stream-selector gate, a canned error-tail), with no write surface. Each target in the config names its own platform, so one config can span the whole stack. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.observability-aiops/ (relocatable via OBSERVABILITY_AIOPS_HOME).\n  Credentials: the Grafana service-account/API token (required) or the Prometheus bearer token (optional; self-hosted Prometheus is often unauthenticated) is stored ENCRYPTED in ~/.observability-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'observability-aiops init' to onboard (it asks for the platform), or 'observability-aiops secret set <target>' to add one. The store is unlocked by a master password from OBSERVABILITY_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var OBSERVABILITY_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback with a deprecation warning (migrate with 'observability-aiops secret migrate'). The token is sent as an Authorization: Bearer header and held only in memory; secrets are never logged or echoed.\n  PromQL is used only through read endpoints (/api/v1/query, /query_range) — there is no write query path. State-changing operations pass through the @governed_tool decorator (budget guard + audit + risk-tier tagging). The destructive write (delete_dashboard) is high-risk with dry_run and captures the full prior dashboard model BEFORE deleting; reversible writes (update_dashboard, create_silence) capture the real fetched before-state and record an inverse undo descriptor. Silences are TIME-BOXED (create_silence requires a positive duration).\n  Webhooks: none — no outbound network calls beyond the configured Prometheus / Alertmanager / Grafana endpoints.\n  SSL: verify_ssl defaults to true; disable for self-signed lab certs.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Verification status: the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (reads, the three metric RCAs, silence + dashboard governed writes, and undo replay); the Loki surface is mock-only so far. Prometheus, Grafana and Loki are free/open-source (docker run prom/prometheus, grafana/grafana, grafana/loki) so a live 'doctor' check is easy. docs/VERIFICATION.md records what was and was not covered.\n---\n\n# Observability AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the Prometheus or Grafana projects, Grafana Labs, or the CNCF.** Prometheus, Alertmanager and Grafana are trademarks of their respective owners. Source at [github.com/AIops-tools/Observability-AIops](https://github.com/AIops-tools/Observability-AIops) under the MIT license.\n\nGoverned self-hosted observability operations — **39 MCP tools** across\n**Prometheus** (HTTP API + PromQL), **Alertmanager** (alerts + silences),\n**Grafana** (dashboards, datasources, folders), and **Grafana Loki** (bounded\nLogQL log reads + log RCA, read-only), every one wrapped with the bundled\n`@governed_tool` harness: a local unified audit log under\n`~/.observability-aiops/`, token/runaway budget guard, undo-token\nrecording, and descriptive risk-tier labels. One config can span the whole\nstack. Bearer tokens are stored **encrypted** (`~/.observability-aiops/secrets.enc`,\nFernet + scrypt) — never plaintext on disk.\n\nThis is the **self-hosted-observability** complement to enterprise monitoring\nsuites: it speaks the open Prometheus/Grafana APIs an SRE actually runs.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`observability_aiops.governance`) — no external skill-family dependency.\n> Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been\n> exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki\n> surface has not yet been exercised live (see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **Metrics** | Prometheus | instant_query, range_query, label_values, series_metadata | 4 | read |\n| **Targets & status** | Prometheus | list_targets, target_scrape_health, dropped_targets, prometheus_config_status, prometheus_tsdb_status | 5 | read |\n| **Rules** | Prometheus | list_rules, rule_health | 2 | read |\n| **Alerts** | Prometheus/Alertmanager | firing_alerts, pending_alerts, alertmanager_alerts, list_silences | 4 | read |\n| **Grafana** | Grafana | list_dashboards, get_dashboard, list_datasources, datasource_health, list_folders | 5 | read |\n| **Loki** | Loki | loki_labels, loki_label_values, loki_query, loki_tail_errors | 4 | read |\n| **Overview + analyses** | all | observability_overview + firing_alert_rca, target_scrape_health_analysis, alert_noise_and_flap_analysis | 4 | read |\n| **Log analyses + cross-signal** | Loki (+ Prometheus) | log_error_burst_rca, log_volume_analysis, alert_log_context | 3 | read |\n| **Writes** | Alertmanager/Grafana/Prometheus | create_silence, expire_silence (med) · create_annotation (med) · update_dashboard (med) · delete_dashboard (**high**) · reload_prometheus_config (med) | 6 | write |\n\nThe three metric flagship analyses are transparent heuristics that report their\nnumbers: `firing_alert_rca` joins each firing alert to its rule expression and\nmaps it to a cause + action; `target_scrape_health_analysis` ranks down/erroring\nscrape targets and classifies each `lastError`; `alert_noise_and_flap_analysis`\nfinds noisy/duplicate alerts and recommends a dedup/rollup. The two **log**\nanalyses mirror this: `log_error_burst_rca` compares per-stream error counts\nagainst a baseline window and classifies each burst (new signature / volume spike\n/ single-instance); `log_volume_analysis` ranks the highest-volume streams and\nwarns on high-cardinality (high-churn) labels. `alert_log_context` bridges the two\nsignals — it maps a firing Prometheus alert's labels to a Loki stream selector and\npulls the correlated logs. **Loki is read-only** (no safe write surface).\n\n## Quick Install\n\n```bash\nuv tool install observability-aiops\nobservability-aiops init       # wizard: pick platform (prometheus/grafana) + encrypted token\nobservability-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/observability-aiops\nopenclaw skills info observability-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a snapshot (`overview` / `observability_overview`): firing-alert count,\n  scrape targets up/down, rules erroring (Prometheus) or dashboard/datasource\n  counts (Grafana)\n- Run PromQL (`instant_query` / `range_query`), enumerate `label_values` or\n  `series_metadata`\n- Check scrape health (`target_scrape_health`, `dropped_targets`) and rule health\n  (`rule_health`, `list_rules`)\n- Triage alerts: `firing_alerts` / `pending_alerts`, the Alertmanager view\n  (`alertmanager_alerts`, `list_silences`), then `firing_alert_rca` to root-cause\n- Reduce alert noise (`alert_noise_and_flap_analysis`) → group_by / inhibition /\n  longer `for`\n- Grafana: `list_dashboards`, `get_dashboard`, `list_datasources`,\n  `datasource_health`, `list_folders`\n- Loki logs: enumerate `loki_labels` / `loki_label_values`, run a bounded\n  `loki_query` (LogQL, stream selector required), `loki_tail_errors` for a\n  selector; then `log_error_burst_rca` to root-cause an error burst and\n  `log_volume_analysis` for volume/cardinality; `alert_log_context` to pull the\n  logs behind a firing alert\n- Governed writes: silence an alert (`create_silence`, time-boxed), annotate an\n  event (`create_annotation`), update/delete a dashboard (`dry_run` first for\n  either), or hot-reload Prometheus (`reload_prometheus_config`)\n\n**Do NOT use when** the target is not a Prometheus/Grafana observability stack —\nroute hypervisor, storage, backup, container-orchestrator, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill. Hosted/SaaS\nmonitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Prometheus / Alertmanager / Grafana observability ops | **observability-aiops** (this skill) |\n| A different platform (hypervisor, storage, backup, orchestrator, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Hosted/SaaS monitoring (Datadog, New Relic, enterprise NMS) | out of scope for this tool |\n\n## Common Workflows\n\n> The CLI covers the reads and the three RCAs (`alert`, `query`, `logs`,\n> `overview`); the guarded **writes** (silences, annotations, dashboards,\n> config reload) are MCP tools — those steps name the tool rather than a CLI\n> command.\n\n### \"Pager went off\" — root-cause the firing alerts and time-box the noise\n\n1. `observability-aiops overview` → one-shot stack picture: firing counts, target\n   health, rule health — is this one alert or the whole stack?\n2. `observability-aiops alert firing` → what is firing right now, grouped by\n   severity\n3. `observability-aiops alert rca` → each firing alert joined to its rule\n   expression with a likely cause and a recommended action (advisory heuristic —\n   verify it, do not act on it blind)\n4. `observability-aiops query instant '<the rule expr>'` → evaluate the alert's\n   own expression yourself and confirm the RCA's reading of it\n5. `observability-aiops query range '<expr>' --start <rfc3339> --end <rfc3339> --step 60s`\n   → see when it crossed the threshold, which usually names the change that\n   caused it\n6. Time-box the noise while you fix the cause: the `create_silence` MCP tool on a\n   specific matcher (a positive duration is **required** — silences cannot be\n   open-ended), then `observability-aiops alert silences` to confirm it landed\n7. **Failure branch**: if the silence was too broad, `expire_silence` ends it\n   immediately, or `observability-aiops undo apply <id>` replays the recorded\n   inverse (`create_silence`'s undo is expire). If `alert rca` returns nothing\n   while alerts are visibly firing, the alerts are coming from Alertmanager\n   without a matching Prometheus rule — check `alertmanager_alerts` and\n   `list_rules` rather than assuming the RCA is broken.\n\n### Investigate a scrape gap (\"metrics went missing\")\n\n1. `observability-aiops overview` → up/down target counts at a glance\n2. `target_scrape_health` → the unhealthy targets with their raw `lastError`\n3. `target_scrape_health_analysis` → down targets ranked, each `lastError`\n   classified (connection refused / timeout / auth / DNS / TLS) with a concrete fix\n4. `dropped_targets` → if a target is missing **entirely** rather than down, it\n   was relabeled away; this is where that shows up\n5. `observability-aiops query instant 'up{job=\"<job>\"}'` → confirm the gap in the\n   metric itself, not just in the target page\n6. After fixing scrape config, `reload_prometheus_config` (a governed write) →\n   then re-run `target_scrape_health` to confirm the target came back\n7. **Failure branch**: if `reload_prometheus_config` succeeds but the target is\n   still down, the config on disk was not what you thought — check\n   `prometheus_config_status` for what Prometheus actually loaded. A reload with\n   a broken config is rejected by Prometheus and leaves the old config running,\n   so a failed reload is not an outage.\n\n### Tame a noisy / flapping alert\n\n1. `observability-aiops alert firing` → the volume of what is firing\n2. `alert_noise_and_flap_analysis` → alertnames with many instances or exact\n   duplicates, each with a `group_by` / inhibition / longer-`for` recommendation\n3. `list_rules` and `rule_health` → read the offending rule's current `for`\n   duration and confirm it is evaluating cleanly\n4. `observability-aiops query range '<rule expr>' --start <rfc3339> --end <rfc3339> --step 60s`\n   → see the flapping in the data and pick a `for` window that actually covers it\n5. `create_silence` for a time-boxed quiet period while the rule change ships;\n   `observability-aiops alert silences` to confirm\n6. **Failure branch**: silencing is a stopgap, not a fix — if the silence expires\n   and the flapping returns, the rule threshold or `for` window is still wrong.\n   Use `observability-aiops undo list` to see exactly which silences this tool\n   created, so no stale silence quietly hides a real outage.\n\n### Root-cause a log error burst (Loki, read-only)\n\n1. `alert_log_context <alertname>` → the firing alert's labels mapped to a Loki\n   stream selector plus the correlated error logs (or start from a selector directly)\n2. `observability-aiops logs errors '{app=\"api\"}' --hours 2 --limit 200` → tail\n   the error-level lines for that stream\n3. `log_error_burst_rca <selector>` → per-stream error counts against a baseline\n   window, each burst classified (new signature / volume spike / single instance)\n4. `observability-aiops logs query '{app=\"api\"} |= \"timeout\"' --hours 2` → confirm\n   the specific signature the RCA named\n5. `log_volume_analysis <selector>` → the highest-volume streams and any\n   high-cardinality label driving a stream/index explosion\n6. **Failure branch**: Loki here is **read-only and bounded** — queries require a\n   stream selector and are capped by lookback and line count. A query rejected\n   for a missing selector is the guard working, not a bug: narrow it with\n   `observability-aiops logs labels` first. There is no write surface for Loki,\n   so remediation happens in the emitting service, not through this tool.\n\n### Safely change or retire a Grafana dashboard (reversible)\n\n1. `list_dashboards` / `list_folders` → locate the dashboard and its folder\n2. `get_dashboard <uid>` → confirm this is the right dashboard before touching it\n3. `update_dashboard` with `dry_run=True` → preview; then for real — it fetches\n   and stashes the **prior model** and records a restore undo\n4. To retire one: `delete_dashboard <uid>` with `dry_run=True` first. Delete is\n   `high` risk — the prior model is captured **before** the delete so the undo\n   can recreate it; set `OBSERVABILITY_AUDIT_APPROVED_BY` (and\n   `OBSERVABILITY_AUDIT_RATIONALE`) if you want that recorded on the audit row\n5. `create_annotation` → mark the change on the timeline so the next responder\n   can correlate a metric shift with this edit\n6. **Failure branch**: wrong dashboard or a bad edit — `observability-aiops undo list`\n   then `observability-aiops undo apply <id>` restores the captured prior model\n   (or recreates a deleted dashboard from it). If the write fails outright, that\n   is the connecting account's permissions (this tool does not gate it) — check\n   the token's role before assuming `observability-aiops doctor` connectivity is\n   at fault.\n\n## Governance & Safety\n\nThe skill delivers reads and writes and records them; it does **not** decide whether a write is\npermitted. That is your agent's judgement, or the permission of the account you connect it with\n(give it a Grafana token with only Viewer scope, and a Prometheus/Alertmanager reached without the\nadmin/write API — writes then fail at the server). There is no read-only switch, policy file, or\napproval gate.\n\n- **Audit is the guarantee, and it is not bypassable.** Every operation — MCP and CLI alike — is\n  logged to `~/.observability-aiops/audit.db` (relocatable via `OBSERVABILITY_AIOPS_HOME`): params,\n  result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.\n- `OBSERVABILITY_AUDIT_APPROVED_BY` / `OBSERVABILITY_AUDIT_RATIONALE` are optional annotations\n  recorded on the audit row (who/why); they are never required and never block.\n- **Runaway guard** — a safety backstop, not authorization: the same call looped in a tight window\n  trips a circuit breaker. Disable with `OBSERVABILITY_RUNAWAY_MAX=0`.\n- Writes support `--dry-run` / `dry_run=True` and double confirmation at the CLI.\n- Silences are **time-boxed** (require a positive duration). Reversible writes\n  capture the real fetched before-state and record an inverse descriptor\n  (create_silence→expire, update/delete dashboard→restore/recreate).\n\n## References\n\n- `references/capabilities.md` — full tool + platform + API-path reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, credentials, and connectivity\n- `references/agent-guardrails.md` — running this with a smaller / local model:\n  what the harness enforces for you, and a ready-made system prompt for the rest\n\nFile v0.10.3:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"observability-aiops\",\n  \"version\": \"0.10.3\",\n  \"publishedAt\": 1789452753053\n}\n\nFile v0.10.3:references/agent-guardrails.md\n\n# Agent guardrails — running observability-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the API did not return comes back as `null`, never as `\"\"`. An alert with no `severity` label, a scrape target that has never errored, a recording rule with no alert `state`, a silence with no `comment` — all report `null`, distinguishable from a genuinely empty value. |\n| \"Tell me if the output was cut off\" | Bounded reads (`loki_query`, `loki_tail_errors`, `loki_labels`, `loki_label_values`, `instant_query`, `range_query`, `label_values`, `series_metadata`, `undo_list`) return `{\"returned\": N, \"limit\": L, \"truncated\": true/false}` alongside the rows. For the Loki reads truncation is **measured** — one line beyond the limit is requested — not guessed from a length coincidence. |\n| \"Preserve the ordering / tell me what's most urgent\" | The analysis tools already return worst-first: `firing_alert_rca` ranks by severity, `target_scrape_health_analysis` puts down targets before slow ones, `alert_noise_and_flap_analysis` sorts by instance count. Priority is the list order, and each entry carries the measured number it was ranked on. |\n| \"Confirm before anything destructive\" | Write tools take `dry_run` and the CLI adds double confirmation. |\n| \"Log what you did\" | Every call is audited to `~/.observability-aiops/audit.db` regardless of what the model says it did, and reversible writes record an undo token (`undo_list` / `undo_apply`). |\n| \"Don't hammer the same call in a loop\" | The runaway guard trips a circuit breaker on tight poll/retry loops — a safety backstop, not authorization. |\n\nAuthorization is not this tool's job. Whether a write is allowed to happen is\ndecided by the account you connect it with, or by your agent's own judgement —\nnot by this harness. See \"Recommended setup for a local model\" below for how to\nenforce read-only at the account instead of in a prompt.\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a self-hosted observability stack (Prometheus, Alertmanager,\nGrafana, Loki) through the observability-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current state of the stack, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit (or a narrower\n  selector) instead of treating the partial result as complete.\n- A null field means the API did not return that value. Report it as \"not\n  available\" — never infer it. A missing \"severity\" label is not \"info\".\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  alert names, severities, label values, or health strings.\n- When an analysis tool returns ranked findings, work in the order given and\n  cite the measured number the ranking is based on.\n\nQUERIES\n- PromQL goes to instant_query / range_query; LogQL goes to loki_query. They are\n  different languages — do not send one to the other.\n- A LogQL query MUST carry a stream selector (e.g. '{app=\"api\"}'). A query\n  without one is rejected by the tool, not silently widened.\n- Use label_values / loki_labels / loki_label_values to discover real label\n  names and values before writing a selector. Do not guess a job, instance, or\n  app name.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, saturation, or regression unless a tool result\n  supports it.\n- Do not add generic advice that does not follow from the tool output.\n- Keep the identifiers straight: an alertname is not a label value; a silence ID\n  is not a dashboard UID; a \"job\" is a scrape config name while an \"instance\" is\n  a single scraped endpoint.\n```\n\n## Recommended setup for a local model\n\nThere is no read-only switch to set — this tool does not decide whether a\nwrite is permitted. If you want the connection to be read-only until you trust\nthe setup, enforce it at the account: give it a Grafana token with only Viewer\nscope, and a Prometheus/Alertmanager reached without the admin/write API. Any\nwrite attempt then fails at the server, which is the place that actually owns\nthe permission.\n\n```bash\nobservability-aiops doctor\n```\n\nWhen you are ready to allow writes (silences, annotations, dashboards), connect\nwith a token that has write scope, and optionally name yourself on the audit\nrow — it is an annotation, not a gate:\n\n```bash\nexport OBSERVABILITY_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport OBSERVABILITY_AUDIT_RATIONALE=\"silencing NodeDiskFilling during the 2026-07-20 disk swap\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the analysis tools —\n  `firing_alert_rca`, `target_scrape_health_analysis`,\n  `alert_noise_and_flap_analysis`, `log_error_burst_rca`, and\n  `alert_log_context` do the multi-step correlation inside one call, so the\n  model does not have to chain reads and keep alertnames, jobs, and selectors\n  straight across turns.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions, scope PromQL and LogQL with real label matchers, and use `--limit`\n  deliberately rather than pulling whole label sets or metric-name lists.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\n## Verification status\n\nThese tools have been exercised against a real stack — Prometheus 3.x,\nAlertmanager, and Grafana 13 — covering firing-alert and scrape-target RCA,\ngoverned silence and dashboard writes, and undo replay. The behaviours described\nabove are observed, not only unit-tested.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Observability-AIops](https://github.com/AIops-tools/Observability-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.3:references/capabilities.md\n\n# observability-aiops capability matrix\n\n> **39 MCP tools** (32 read, 7 write) across Prometheus\n> (HTTP API + PromQL, default port 9090, optional bearer token), a companion\n> Alertmanager (`/api/v2`, port 9093), Grafana (HTTP API, port 3000, required\n> bearer token), and Grafana Loki (HTTP API, port 3100, optional bearer/basic\n> auth, optional multi-tenant `X-Scope-OrgID`). Loki is **read-only**. The\n> Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md).\n\n## Metrics — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `instant_query` | `/api/v1/query` | PromQL evaluated at one instant (samples: metric + value + timestamp) |\n| `range_query` | `/api/v1/query_range` | PromQL over a time range (per-series point arrays) |\n| `label_values` | `/api/v1/label/<name>/values` | distinct values of a label (default `__name__` = all metric names) |\n| `series_metadata` | `/api/v1/series` | series (label-set) metadata for a selector |\n\n## Targets & status — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_targets` | `/api/v1/targets` | active scrape targets (job, instance, health, lastError), optional up/down filter |\n| `target_scrape_health` | `/api/v1/targets` | up/down summary + the unhealthy targets |\n| `dropped_targets` | `/api/v1/targets` | targets discovered but dropped by relabeling |\n| `prometheus_config_status` | `/api/v1/status/config` | running-config fingerprint (sha256) + size — never the raw YAML/secrets |\n| `prometheus_tsdb_status` | `/api/v1/status/tsdb` | TSDB head cardinality + top metrics by series count |\n\n## Rules — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_rules` | `/api/v1/rules` | recording + alerting rules (name, type, expr, health), optional type filter |\n| `rule_health` | `/api/v1/rules` | rule-evaluation health summary + erroring rules |\n\n## Alerts — Prometheus + Alertmanager (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `firing_alerts` | `/api/v1/alerts` | firing Prometheus rule alerts, grouped by severity |\n| `pending_alerts` | `/api/v1/alerts` | pending (not-yet-firing) rule alerts |\n| `alertmanager_alerts` | AM `/api/v2/alerts` | alerts as Alertmanager sees them (post grouping/silence/inhibit) |\n| `list_silences` | AM `/api/v2/silences` | silences (active, pending, expired) with matchers |\n\n## Grafana (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_dashboards` | `/api/search?type=dash-db` | dashboards (uid, title, folder, tags), optional title query |\n| `get_dashboard` | `/api/dashboards/uid/{uid}` | one dashboard's summary (title, version, panel + tag counts) |\n| `list_datasources` | `/api/datasources` | datasources (id, uid, name, type, default flag) |\n| `datasource_health` | `/api/datasources/{id}/health` | one datasource's health (status, message) |\n| `list_folders` | `/api/folders` | Grafana folders |\n\n## Loki — logs (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `loki_labels` | `/loki/api/v1/labels` | distinct label names in the lookback window |\n| `loki_label_values` | `/loki/api/v1/label/<name>/values` | distinct values of one label (name percent-encoded) |\n| `loki_query` | `/loki/api/v1/query_range` | bounded LogQL passthrough — **requires a stream selector**; lookback capped at `MAX_LOOKBACK_HOURS=24`, lines clamped to `MAX_LINE_LIMIT=1000` (default 100) |\n| `loki_tail_errors` | `/loki/api/v1/query_range` | canned error-level read for a selector (line-filter on `(?i)(error\\|fatal\\|panic\\|exception\\|traceback\\|stacktrace)`) |\n\nBounding gate: a LogQL query with **no `{…}` stream selector**, an empty query,\nor a lookback beyond the cap is rejected up front with a teaching error — no\nunbounded scan is ever issued. Label values interpolated into a selector are\nbackslash-escaped; label names in a path are percent-encoded. Loki auth is\noptional (bearer or, per target, `basic` with a `user:password` secret) and a\nmulti-tenant `X-Scope-OrgID` header is sent when the target sets `org_id`.\n\n## Overview & flagship analyses (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `observability_overview` | platform-aware | Prometheus: firing count + targets up/down + rules erroring; Grafana: dashboard/datasource/folder counts; Loki: label-name count |\n| `firing_alert_rca` | firing alerts + alerting rules | each firing alert joined to its rule expr, ranked by severity, mapped to a likely **cause + action** |\n| `target_scrape_health_analysis` | active targets | down/erroring scrapes ranked, each `lastError` classified (refused/timeout/auth/DNS/TLS) with a fix |\n| `alert_noise_and_flap_analysis` | alert instances | alertnames with many instances / exact duplicates flagged with a group_by / inhibition / longer-`for` recommendation |\n\n## Loki — log analyses & cross-signal (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `log_error_burst_rca` | selector + window (pulls current + baseline error streams) | per-stream error counts vs a baseline window; each burst classified **new_signature** (baseline 0), **volume_spike** (>= `burst_ratio`×baseline), or **single_instance** (localized to one pod/instance) with a cause + action + sample lines |\n| `log_volume_analysis` | selector + window (pulls streams + `index/stats`) | top streams by line volume, high-cardinality (high-churn) label warnings, and a retention hint from total ingest bytes |\n| `alert_log_context` | firing alertname (Prometheus) + Loki target | maps the alert's labels → a Loki stream selector (intersect with namespace/job/service/app/container/pod/instance/component, first 4 in priority order; values escaped) and returns the correlated error streams. **Best-effort**: only labels the alert and Loki share will match |\n\n## Undo (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `undo_list` | local undo store (`limit`) | recorded, not-yet-applied reversible writes: undoId, ts, originalTool, inverseTool, note |\n\n## Writes (governed)\n\n| Tool | Risk | API path | Notes |\n|------|------|----------|-------|\n| `create_silence` | **med** | AM `POST /api/v2/silences` | **time-boxed** (requires minutes > 0); returns silenceId; undo → `expire_silence` |\n| `expire_silence` | **med** | AM `DELETE /api/v2/silence/{id}` | inverse of create_silence |\n| `create_annotation` | **medium** | `POST /api/annotations` | Grafana event marker |\n| `update_dashboard` | **med** | `POST /api/dashboards/db` | GETs the prior model first → captures it for a restore undo |\n| `delete_dashboard` | **HIGH** | `DELETE /api/dashboards/uid/{uid}` | `dry_run`; captures prior model **BEFORE** delete; undo → recreate |\n| `reload_prometheus_config` | **med** | `POST /-/reload` | records the pre-reload config hash; no undo (re-apply the prior config file) |\n| `undo_apply` | **med** | dispatches the recorded inverse tool | executes a recorded inverse; the inverse runs through its own governed tool (its real risk tier applies); single-use token; supports `dry_run` |\n\n**No Loki writes.** Loki exposes no safe operational write surface here (no\nsilence/annotation analogue), so this tool ships Loki as read-only by design.\n\n## Out of scope (by design)\n\n- **Hosted/SaaS monitoring** — Datadog, New Relic, and enterprise NMS (only\n  self-hosted Prometheus + Grafana + Loki here)\n- **Loki writes / ingestion / deletes** — read-only LogQL only (no push, no\n  delete-series, no ruler/config writes)\n- **Creating/editing Prometheus rules or scrape config**, and provisioning\n  Grafana datasources/dashboards from scratch (beyond update/delete of an existing\n  dashboard)\n- **Long-term-storage query fan-out** (Thanos/Cortex/Mimir) — the single\n  Prometheus HTTP API only\n\nFile v0.10.3:references/cli-reference.md\n\n# observability-aiops CLI reference\n\n> Covers Prometheus (HTTP API + PromQL), a companion Alertmanager, Grafana\n> (HTTP API), and Grafana Loki (LogQL, read-only). The Prometheus/Alertmanager/\n> Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack;\n> the Loki surface has not (see docs/VERIFICATION.md). The CLI is a convenience\n> subset — the full 39-tool surface is via the MCP server\n> (`observability-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nobservability-aiops init                      # interactive wizard (asks for the platform: prometheus/grafana/loki)\nobservability-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   Prometheus: /api/v1/status/buildinfo · Grafana: /api/health\n                                           #   Loki: /ready + /loki/api/v1/status/buildinfo\nobservability-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.observability-aiops/secrets.enc)\n\n```bash\nobservability-aiops secret set <target> [--value <token>]   # store bearer token (hidden prompt if no --value)\nobservability-aiops secret list                             # names only — secrets never shown\nobservability-aiops secret rm <target>\nobservability-aiops secret migrate                          # import legacy plaintext env (OBSERVABILITY_<TARGET>_TOKEN)\nobservability-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nobservability-aiops overview [--target <t>]   # snapshot: firing alerts + targets up/down + rules erroring (Prometheus)\n                                           #   or dashboard/datasource/folder counts (Grafana) / label-name count (Loki)\n```\n\n## Query (Prometheus PromQL)\n\n```bash\nobservability-aiops query instant 'up'                      # PromQL instant query\nobservability-aiops query range 'rate(x[5m])' --start ... --end ... [--step 60s]\nobservability-aiops query labels [__name__]                 # distinct label values (default = all metric names)\n```\n\n## Logs (Grafana Loki, read-only, bounded)\n\n```bash\nobservability-aiops logs labels [--hours 1] [--target <t>]              # distinct Loki label names in the window\nobservability-aiops logs query '{app=\"api\"} |= \"error\"' [--hours 1] [--limit 100]   # bounded LogQL (stream selector required)\nobservability-aiops logs errors '{app=\"api\"}' [--hours 1] [--limit 100]  # canned error-level tail for a selector\n```\n\n## Alerts\n\n```bash\nobservability-aiops alert firing [--target <t>]             # firing Prometheus rule alerts, by severity\nobservability-aiops alert silences [--target <t>]           # Alertmanager silences\nobservability-aiops alert rca [--target <t>]                # root-cause firing alerts (join to rule expr → cause+action)\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `query`, `logs`, and `alert` are the CLI subset; the remaining\n  metrics, targets, rules, Grafana, Loki analyses (log_error_burst_rca,\n  log_volume_analysis, alert_log_context), and governed-write tools\n  (create/expire silence, create annotation, update/delete dashboard, reload\n  config) are exposed through the MCP server. High-risk MCP writes honour `OBSERVABILITY_AUDIT_APPROVED_BY` /\n  `OBSERVABILITY_AUDIT_RATIONALE` and support dry-run.\n\nFile v0.10.3:references/setup-guide.md\n\n# observability-aiops setup & security guide\n\n> The Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md). **Prometheus,\n> Grafana, and Loki are all free/open-source and trivial to stand up in a lab\n> (`docker run prom/prometheus`, `grafana/grafana`, `grafana/loki`), so a live\n> `doctor` check is easy.**\n\n## 1. Install\n\n```bash\nuv tool install observability-aiops\n```\n\n## 2. Get a credential\n\n- **Prometheus** — a bearer token is **optional**; many self-hosted deployments\n  are unauthenticated. observability-aiops talks to the HTTP API on port **9090**\n  and a companion Alertmanager on **9093** (`/api/v2`).\n- **Grafana** — a **service-account token** (Administration → Service accounts →\n  Add token) or legacy API key is **required**. Grafana's HTTP API is on port\n  **3000**.\n- **Loki** — auth is **optional** (bearer token, or per-target `basic` auth where\n  the stored secret is `user:password`). The HTTP API is on port **3100**. For a\n  multi-tenant deployment set `org_id` to send the `X-Scope-OrgID` header. Loki is\n  **read-only** here.\n\n## 3. Onboard\n\n```bash\nobservability-aiops init\n```\n\nThe wizard asks, per target, for the **platform** (`prometheus` / `grafana` /\n`loki`), the **host**, the **scheme** (`http` / `https`), the **port** (defaults\n9090 for Prometheus, 3000 for Grafana, 3100 for Loki), an optional **Alertmanager\nURL** (Prometheus only), the **auth type** and **org id** (Loki only), and the\n**token** — required for Grafana, optional for Prometheus/Loki. Non-secret\nconnection details go to `~/.observability-aiops/config.yaml`; the token is stored\n**encrypted** into `~/.observability-aiops/secrets.enc`. Example config (one config\ncan span the whole stack):\n\n```yaml\ntargets:\n  - name: prod-prom\n    platform: prometheus\n    host: 10.0.0.20\n    scheme: http\n    port: 9090\n    alertmanager_url: http://10.0.0.20:9093   # optional; blank assumes host:9093\n  - name: prod-grafana\n    platform: grafana\n    host: 10.0.0.30\n    scheme: https\n    port: 3000\n    verify_ssl: true\n  - name: prod-loki\n    platform: loki\n    host: 10.0.0.40\n    scheme: http\n    port: 3100\n    auth_type: bearer        # or 'basic' (secret is user:password)\n    org_id: team-a           # optional; sent as X-Scope-OrgID (multi-tenant)\n```\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nExport the master password so the encrypted store can be unlocked without a\nprompt:\n\n```bash\nexport OBSERVABILITY_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## Credential security\n\n- The token is **never** written to disk in plaintext. It lives only in\n  `~/.observability-aiops/secrets.enc`, encrypted with Fernet (AES-128-CBC +\n  HMAC), the key derived from your master password via scrypt. Only a per-store\n  random salt and the ciphertext are on disk (chmod 600); the master password\n  itself is never stored.\n- A legacy plaintext env var `OBSERVABILITY_<TARGET_NAME_UPPER>_TOKEN` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `observability-aiops secret migrate` (it imports then renames the old `.env`).\n- The token is sent as an `Authorization: Bearer` header at request time and held\n  only in memory; it is never logged or echoed. Exception text and tracebacks are\n  scrubbed of secret-shaped strings before being written to the audit log.\n\n## Governance harness state\n\nState lives under `~/.observability-aiops/` (relocate with `OBSERVABILITY_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with risk tier and any operator-supplied\n  approver/rationale (optional annotations, never required)\n- `undo.db` — inverse descriptors for reversible writes (create_silence→expire,\n  update/delete dashboard→restore/recreate)\n- budget / runaway guard — caps cumulative tool calls and wall-time; trips on\n  tight poll/retry loops\n\n## Governed writes\n\n- **High-risk** op (`delete_dashboard`) supports `dry_run` and captures the full\n  prior dashboard model **before** deleting so the recorded undo can recreate\n  it. Optionally set `OBSERVABILITY_AUDIT_APPROVED_BY` and\n  `OBSERVABILITY_AUDIT_RATIONALE` to annotate the audit row — neither is\n  required, and the write runs either way.\n- **Reversible** writes capture the real fetched before-state:\n  `update_dashboard` (restore prior model), `create_silence` (expire the created\n  silence). `reload_prometheus_config` records the pre-reload config hash.\n- **Time-boxed** ops require a positive duration: `create_silence` (in minutes).\n  This prevents forgotten, indefinite silences.\n\n## Verify\n\n```bash\nobservability-aiops doctor\n```\n\n`doctor` is platform-aware: it checks the config file, the encrypted store and its\npermissions, that a token is present where required, and (unless `--skip-auth`)\nconnectivity — `/api/v1/status/buildinfo` for Prometheus targets, `/api/health`\nfor Grafana targets, and `/ready` + `/loki/api/v1/status/buildinfo` for Loki\ntargets.\n\n## Loki query bounding (safety)\n\nLoki reads are deliberately bounded so an agent can't ask for \"all logs, forever\":\n\n- Every `loki_query` must carry a `{…}` **stream selector** — an unbounded query\n  with no selector (or an empty query) is rejected with a teaching error.\n- Lookback is capped at **24h** (`MAX_LOOKBACK_HOURS`); a longer window is refused.\n- Returned lines are clamped to **1000** (`MAX_LINE_LIMIT`, default 100).\n- `loki_tail_errors` wraps a selector with a canned case-insensitive error filter.\n- Label values interpolated into a selector are backslash-escaped and label names\n  in a path segment are percent-encoded, so a hostile label value can't break out\n  of the LogQL string or rewrite the request path.\n\nFile v0.10.3:skill-card.md\n\n## Description:\n\nOperates self-hosted Prometheus, Alertmanager, Grafana, and Grafana Loki stacks with governed reads, root-cause analyses, bounded log queries, and audited write actions.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, SREs, and operations engineers use this skill to inspect and operate self-hosted Prometheus, Alertmanager, Grafana, and Loki environments, including alert triage, scrape-health investigation, dashboard work, and bounded log analysis.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The release installs an unpinned external executable.\n\nMitigation: Install only from a trusted publisher and pin or independently verify the executable before use.\n\nRisk: Configured credentials can make real changes to Grafana, Alertmanager, or Prometheus.\n\nMitigation: Use least-privilege accounts and enforce read-only behavior through server-side permissions when write access is not intended.\n\nRisk: Authenticated observability traffic may expose operational data or credentials if certificate verification is disabled.\n\nMitigation: Use HTTPS with certificate verification enabled and reserve verify_ssl: false for isolated lab environments only.\n\nRisk: Dashboard deletion and other write tools can alter production observability state.\n\nMitigation: Use dry-run paths where available, review audited actions, and rely on undo records for reversible write operations.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/observability-aiops)\n- [Project homepage](https://github.com/AIops-tools/Observability-AIops)\n- [Agent guardrails](references/agent-guardrails.md)\n- [Capability matrix](references/capabilities.md)\n- [CLI reference](references/cli-reference.md)\n- [Setup and security guide](references/setup-guide.md)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Shell commands, Configuration, Guidance]\n\n**Output Format:** [Markdown and structured text with inline shell commands and operational findings]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include bounded query results, RCA summaries, audit and undo guidance, and dry-run recommendations for write actions.]\n\n## Skill Version(s):\n\n0.10.3 (source: ClawHub release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.2: 7 files, 21133 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2495b), SKILL.md (20269b), _meta.json (139b)\n\nFile v0.10.2:SKILL.md\n\n---\nname: observability-aiops\nslug: observability-aiops\ndisplayName: \"Observability AIops\"\nsummary: \"Governed Prometheus + Grafana ops: PromQL, alerts, dashboards, RCA; 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Observability-AIops\ntags: [aiops, mcp, governance, observability]\ndescription: >\n  Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager alerts + silences, Grafana dashboards/datasources/folders, bounded Loki LogQL log reads (labels, query, error-tail), five flagship analyses (firing-alert RCA, target-scrape-health, alert-noise/flap, log-error-burst RCA, log-volume/cardinality) plus an alert->log cross-signal, and guarded writes (create/expire silence, create annotation, update/delete dashboard, reload Prometheus config).\n  Always use this skill for \"Prometheus\", \"PromQL\", \"Alertmanager\", \"Grafana\", \"Loki\", \"LogQL\", \"logs\", \"which targets are down\", \"scrape failing\", \"why is this alert firing\", \"root cause this alert\", \"firing alerts\", \"silence this alert\", \"noisy alerts\", \"alert flapping\", \"recording rule\", \"alerting rule\", \"dashboard\", \"datasource health\", \"reload prometheus config\", \"TSDB cardinality\", \"error burst\", \"log volume\", \"log cardinality\", \"tail errors\" when the context is a self-hosted metrics/logs/observability stack.\n  Do NOT use when the target is something other than a Prometheus/Grafana observability stack (a hypervisor, storage appliance, backup product, container-orchestrator control plane, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n  Governed observability operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (RCAs, governed writes, undo); the Loki surface has not (see docs/VERIFICATION.md).\ninstaller:\n  kind: uv\n  package: observability-aiops\nargument-hint: \"[a PromQL query, an alert/dashboard uid, or describe your observability task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"observability-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"OBSERVABILITY_AIOPS_CONFIG\",\"OBSERVABILITY_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Observability-AIops\",\"emoji\":\"📈\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed observability operations across Prometheus (HTTP API + PromQL, default port 9090, optional bearer token), a companion Alertmanager (/api/v2, default port 9093), Grafana (HTTP API, default port 3000, required bearer token), and Grafana Loki (HTTP API, default port 3100, optional bearer or basic auth, optional multi-tenant X-Scope-OrgID). Loki is READ-ONLY: bounded LogQL reads only (labels, label values, query_range with a hard lookback + line cap and a stream-selector gate, a canned error-tail), with no write surface. Each target in the config names its own platform, so one config can span the whole stack. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.observability-aiops/ (relocatable via OBSERVABILITY_AIOPS_HOME).\n  Credentials: the Grafana service-account/API token (required) or the Prometheus bearer token (optional; self-hosted Prometheus is often unauthenticated) is stored ENCRYPTED in ~/.observability-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'observability-aiops init' to onboard (it asks for the platform), or 'observability-aiops secret set <target>' to add one. The store is unlocked by a master password from OBSERVABILITY_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var OBSERVABILITY_<TARGET_NAME_UPPER>_TOKEN is still honoured as a fallback with a deprecation warning (migrate with 'observability-aiops secret migrate'). The token is sent as an Authorization: Bearer header and held only in memory; secrets are never logged or echoed.\n  PromQL is used only through read endpoints (/api/v1/query, /query_range) — there is no write query path. State-changing operations pass through the @governed_tool decorator (budget guard + audit + risk-tier tagging). The destructive write (delete_dashboard) is high-risk with dry_run and captures the full prior dashboard model BEFORE deleting; reversible writes (update_dashboard, create_silence) capture the real fetched before-state and record an inverse undo descriptor. Silences are TIME-BOXED (create_silence requires a positive duration).\n  Webhooks: none — no outbound network calls beyond the configured Prometheus / Alertmanager / Grafana endpoints.\n  SSL: verify_ssl defaults to true; disable for self-signed lab certs.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Verification status: the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (reads, the three metric RCAs, silence + dashboard governed writes, and undo replay); the Loki surface is mock-only so far. Prometheus, Grafana and Loki are free/open-source (docker run prom/prometheus, grafana/grafana, grafana/loki) so a live 'doctor' check is easy. docs/VERIFICATION.md records what was and was not covered.\n---\n\n# Observability AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by the Prometheus or Grafana projects, Grafana Labs, or the CNCF.** Prometheus, Alertmanager and Grafana are trademarks of their respective owners. Source at [github.com/AIops-tools/Observability-AIops](https://github.com/AIops-tools/Observability-AIops) under the MIT license.\n\nGoverned self-hosted observability operations — **39 MCP tools** across\n**Prometheus** (HTTP API + PromQL), **Alertmanager** (alerts + silences),\n**Grafana** (dashboards, datasources, folders), and **Grafana Loki** (bounded\nLogQL log reads + log RCA, read-only), every one wrapped with the bundled\n`@governed_tool` harness: a local unified audit log under\n`~/.observability-aiops/`, token/runaway budget guard, undo-token\nrecording, and descriptive risk-tier labels. One config can span the whole\nstack. Bearer tokens are stored **encrypted** (`~/.observability-aiops/secrets.enc`,\nFernet + scrypt) — never plaintext on disk.\n\nThis is the **self-hosted-observability** complement to enterprise monitoring\nsuites: it speaks the open Prometheus/Grafana APIs an SRE actually runs.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`observability_aiops.governance`) — no external skill-family dependency.\n> Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been\n> exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki\n> surface has not yet been exercised live (see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **Metrics** | Prometheus | instant_query, range_query, label_values, series_metadata | 4 | read |\n| **Targets & status** | Prometheus | list_targets, target_scrape_health, dropped_targets, prometheus_config_status, prometheus_tsdb_status | 5 | read |\n| **Rules** | Prometheus | list_rules, rule_health | 2 | read |\n| **Alerts** | Prometheus/Alertmanager | firing_alerts, pending_alerts, alertmanager_alerts, list_silences | 4 | read |\n| **Grafana** | Grafana | list_dashboards, get_dashboard, list_datasources, datasource_health, list_folders | 5 | read |\n| **Loki** | Loki | loki_labels, loki_label_values, loki_query, loki_tail_errors | 4 | read |\n| **Overview + analyses** | all | observability_overview + firing_alert_rca, target_scrape_health_analysis, alert_noise_and_flap_analysis | 4 | read |\n| **Log analyses + cross-signal** | Loki (+ Prometheus) | log_error_burst_rca, log_volume_analysis, alert_log_context | 3 | read |\n| **Writes** | Alertmanager/Grafana/Prometheus | create_silence, expire_silence (med) · create_annotation (med) · update_dashboard (med) · delete_dashboard (**high**) · reload_prometheus_config (med) | 6 | write |\n\nThe three metric flagship analyses are transparent heuristics that report their\nnumbers: `firing_alert_rca` joins each firing alert to its rule expression and\nmaps it to a cause + action; `target_scrape_health_analysis` ranks down/erroring\nscrape targets and classifies each `lastError`; `alert_noise_and_flap_analysis`\nfinds noisy/duplicate alerts and recommends a dedup/rollup. The two **log**\nanalyses mirror this: `log_error_burst_rca` compares per-stream error counts\nagainst a baseline window and classifies each burst (new signature / volume spike\n/ single-instance); `log_volume_analysis` ranks the highest-volume streams and\nwarns on high-cardinality (high-churn) labels. `alert_log_context` bridges the two\nsignals — it maps a firing Prometheus alert's labels to a Loki stream selector and\npulls the correlated logs. **Loki is read-only** (no safe write surface).\n\n## Quick Install\n\n```bash\nuv tool install observability-aiops\nobservability-aiops init       # wizard: pick platform (prometheus/grafana) + encrypted token\nobservability-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/observability-aiops\nopenclaw skills info observability-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a snapshot (`overview` / `observability_overview`): firing-alert count,\n  scrape targets up/down, rules erroring (Prometheus) or dashboard/datasource\n  counts (Grafana)\n- Run PromQL (`instant_query` / `range_query`), enumerate `label_values` or\n  `series_metadata`\n- Check scrape health (`target_scrape_health`, `dropped_targets`) and rule health\n  (`rule_health`, `list_rules`)\n- Triage alerts: `firing_alerts` / `pending_alerts`, the Alertmanager view\n  (`alertmanager_alerts`, `list_silences`), then `firing_alert_rca` to root-cause\n- Reduce alert noise (`alert_noise_and_flap_analysis`) → group_by / inhibition /\n  longer `for`\n- Grafana: `list_dashboards`, `get_dashboard`, `list_datasources`,\n  `datasource_health`, `list_folders`\n- Loki logs: enumerate `loki_labels` / `loki_label_values`, run a bounded\n  `loki_query` (LogQL, stream selector required), `loki_tail_errors` for a\n  selector; then `log_error_burst_rca` to root-cause an error burst and\n  `log_volume_analysis` for volume/cardinality; `alert_log_context` to pull the\n  logs behind a firing alert\n- Governed writes: silence an alert (`create_silence`, time-boxed), annotate an\n  event (`create_annotation`), update/delete a dashboard (`dry_run` first for\n  either), or hot-reload Prometheus (`reload_prometheus_config`)\n\n**Do NOT use when** the target is not a Prometheus/Grafana observability stack —\nroute hypervisor, storage, backup, container-orchestrator, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill. Hosted/SaaS\nmonitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Prometheus / Alertmanager / Grafana observability ops | **observability-aiops** (this skill) |\n| A different platform (hypervisor, storage, backup, orchestrator, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Hosted/SaaS monitoring (Datadog, New Relic, enterprise NMS) | out of scope for this tool |\n\n## Common Workflows\n\n> The CLI covers the reads and the three RCAs (`alert`, `query`, `logs`,\n> `overview`); the guarded **writes** (silences, annotations, dashboards,\n> config reload) are MCP tools — those steps name the tool rather than a CLI\n> command.\n\n### \"Pager went off\" — root-cause the firing alerts and time-box the noise\n\n1. `observability-aiops overview` → one-shot stack picture: firing counts, target\n   health, rule health — is this one alert or the whole stack?\n2. `observability-aiops alert firing` → what is firing right now, grouped by\n   severity\n3. `observability-aiops alert rca` → each firing alert joined to its rule\n   expression with a likely cause and a recommended action (advisory heuristic —\n   verify it, do not act on it blind)\n4. `observability-aiops query instant '<the rule expr>'` → evaluate the alert's\n   own expression yourself and confirm the RCA's reading of it\n5. `observability-aiops query range '<expr>' --start <rfc3339> --end <rfc3339> --step 60s`\n   → see when it crossed the threshold, which usually names the change that\n   caused it\n6. Time-box the noise while you fix the cause: the `create_silence` MCP tool on a\n   specific matcher (a positive duration is **required** — silences cannot be\n   open-ended), then `observability-aiops alert silences` to confirm it landed\n7. **Failure branch**: if the silence was too broad, `expire_silence` ends it\n   immediately, or `observability-aiops undo apply <id>` replays the recorded\n   inverse (`create_silence`'s undo is expire). If `alert rca` returns nothing\n   while alerts are visibly firing, the alerts are coming from Alertmanager\n   without a matching Prometheus rule — check `alertmanager_alerts` and\n   `list_rules` rather than assuming the RCA is broken.\n\n### Investigate a scrape gap (\"metrics went missing\")\n\n1. `observability-aiops overview` → up/down target counts at a glance\n2. `target_scrape_health` → the unhealthy targets with their raw `lastError`\n3. `target_scrape_health_analysis` → down targets ranked, each `lastError`\n   classified (connection refused / timeout / auth / DNS / TLS) with a concrete fix\n4. `dropped_targets` → if a target is missing **entirely** rather than down, it\n   was relabeled away; this is where that shows up\n5. `observability-aiops query instant 'up{job=\"<job>\"}'` → confirm the gap in the\n   metric itself, not just in the target page\n6. After fixing scrape config, `reload_prometheus_config` (a governed write) →\n   then re-run `target_scrape_health` to confirm the target came back\n7. **Failure branch**: if `reload_prometheus_config` succeeds but the target is\n   still down, the config on disk was not what you thought — check\n   `prometheus_config_status` for what Prometheus actually loaded. A reload with\n   a broken config is rejected by Prometheus and leaves the old config running,\n   so a failed reload is not an outage.\n\n### Tame a noisy / flapping alert\n\n1. `observability-aiops alert firing` → the volume of what is firing\n2. `alert_noise_and_flap_analysis` → alertnames with many instances or exact\n   duplicates, each with a `group_by` / inhibition / longer-`for` recommendation\n3. `list_rules` and `rule_health` → read the offending rule's current `for`\n   duration and confirm it is evaluating cleanly\n4. `observability-aiops query range '<rule expr>' --start <rfc3339> --end <rfc3339> --step 60s`\n   → see the flapping in the data and pick a `for` window that actually covers it\n5. `create_silence` for a time-boxed quiet period while the rule change ships;\n   `observability-aiops alert silences` to confirm\n6. **Failure branch**: silencing is a stopgap, not a fix — if the silence expires\n   and the flapping returns, the rule threshold or `for` window is still wrong.\n   Use `observability-aiops undo list` to see exactly which silences this tool\n   created, so no stale silence quietly hides a real outage.\n\n### Root-cause a log error burst (Loki, read-only)\n\n1. `alert_log_context <alertname>` → the firing alert's labels mapped to a Loki\n   stream selector plus the correlated error logs (or start from a selector directly)\n2. `observability-aiops logs errors '{app=\"api\"}' --hours 2 --limit 200` → tail\n   the error-level lines for that stream\n3. `log_error_burst_rca <selector>` → per-stream error counts against a baseline\n   window, each burst classified (new signature / volume spike / single instance)\n4. `observability-aiops logs query '{app=\"api\"} |= \"timeout\"' --hours 2` → confirm\n   the specific signature the RCA named\n5. `log_volume_analysis <selector>` → the highest-volume streams and any\n   high-cardinality label driving a stream/index explosion\n6. **Failure branch**: Loki here is **read-only and bounded** — queries require a\n   stream selector and are capped by lookback and line count. A query rejected\n   for a missing selector is the guard working, not a bug: narrow it with\n   `observability-aiops logs labels` first. There is no write surface for Loki,\n   so remediation happens in the emitting service, not through this tool.\n\n### Safely change or retire a Grafana dashboard (reversible)\n\n1. `list_dashboards` / `list_folders` → locate the dashboard and its folder\n2. `get_dashboard <uid>` → confirm this is the right dashboard before touching it\n3. `update_dashboard` with `dry_run=True` → preview; then for real — it fetches\n   and stashes the **prior model** and records a restore undo\n4. To retire one: `delete_dashboard <uid>` with `dry_run=True` first. Delete is\n   `high` risk — the prior model is captured **before** the delete so the undo\n   can recreate it; set `OBSERVABILITY_AUDIT_APPROVED_BY` (and\n   `OBSERVABILITY_AUDIT_RATIONALE`) if you want that recorded on the audit row\n5. `create_annotation` → mark the change on the timeline so the next responder\n   can correlate a metric shift with this edit\n6. **Failure branch**: wrong dashboard or a bad edit — `observability-aiops undo list`\n   then `observability-aiops undo apply <id>` restores the captured prior model\n   (or recreates a deleted dashboard from it). If the write fails outright, that\n   is the connecting account's permissions (this tool does not gate it) — check\n   the token's role before assuming `observability-aiops doctor` connectivity is\n   at fault.\n\n## Governance & Safety\n\nThe skill delivers reads and writes and records them; it does **not** decide whether a write is\npermitted. That is your agent's judgement, or the permission of the account you connect it with\n(give it a Grafana token with only Viewer scope, and a Prometheus/Alertmanager reached without the\nadmin/write API — writes then fail at the server). There is no read-only switch, policy file, or\napproval gate.\n\n- **Audit is the guarantee, and it is not bypassable.** Every operation — MCP and CLI alike — is\n  logged to `~/.observability-aiops/audit.db` (relocatable via `OBSERVABILITY_AIOPS_HOME`): params,\n  result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.\n- `OBSERVABILITY_AUDIT_APPROVED_BY` / `OBSERVABILITY_AUDIT_RATIONALE` are optional annotations\n  recorded on the audit row (who/why); they are never required and never block.\n- **Runaway guard** — a safety backstop, not authorization: the same call looped in a tight window\n  trips a circuit breaker. Disable with `OBSERVABILITY_RUNAWAY_MAX=0`.\n- Writes support `--dry-run` / `dry_run=True` and double confirmation at the CLI.\n- Silences are **time-boxed** (require a positive duration). Reversible writes\n  capture the real fetched before-state and record an inverse descriptor\n  (create_silence→expire, update/delete dashboard→restore/recreate).\n\n## References\n\n- `references/capabilities.md` — full tool + platform + API-path reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, credentials, and connectivity\n- `references/agent-guardrails.md` — running this with a smaller / local model:\n  what the harness enforces for you, and a ready-made system prompt for the rest\n\nFile v0.10.2:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"observability-aiops\",\n  \"version\": \"0.10.2\",\n  \"publishedAt\": 1789223824077\n}\n\nFile v0.10.2:references/agent-guardrails.md\n\n# Agent guardrails — running observability-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the API did not return comes back as `null`, never as `\"\"`. An alert with no `severity` label, a scrape target that has never errored, a recording rule with no alert `state`, a silence with no `comment` — all report `null`, distinguishable from a genuinely empty value. |\n| \"Tell me if the output was cut off\" | Bounded reads (`loki_query`, `loki_tail_errors`, `loki_labels`, `loki_label_values`, `instant_query`, `range_query`, `label_values`, `series_metadata`, `undo_list`) return `{\"returned\": N, \"limit\": L, \"truncated\": true/false}` alongside the rows. For the Loki reads truncation is **measured** — one line beyond the limit is requested — not guessed from a length coincidence. |\n| \"Preserve the ordering / tell me what's most urgent\" | The analysis tools already return worst-first: `firing_alert_rca` ranks by severity, `target_scrape_health_analysis` puts down targets before slow ones, `alert_noise_and_flap_analysis` sorts by instance count. Priority is the list order, and each entry carries the measured number it was ranked on. |\n| \"Confirm before anything destructive\" | Write tools take `dry_run` and the CLI adds double confirmation. |\n| \"Log what you did\" | Every call is audited to `~/.observability-aiops/audit.db` regardless of what the model says it did, and reversible writes record an undo token (`undo_list` / `undo_apply`). |\n| \"Don't hammer the same call in a loop\" | The runaway guard trips a circuit breaker on tight poll/retry loops — a safety backstop, not authorization. |\n\nAuthorization is not this tool's job. Whether a write is allowed to happen is\ndecided by the account you connect it with, or by your agent's own judgement —\nnot by this harness. See \"Recommended setup for a local model\" below for how to\nenforce read-only at the account instead of in a prompt.\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a self-hosted observability stack (Prometheus, Alertmanager,\nGrafana, Loki) through the observability-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current state of the stack, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit (or a narrower\n  selector) instead of treating the partial result as complete.\n- A null field means the API did not return that value. Report it as \"not\n  available\" — never infer it. A missing \"severity\" label is not \"info\".\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  alert names, severities, label values, or health strings.\n- When an analysis tool returns ranked findings, work in the order given and\n  cite the measured number the ranking is based on.\n\nQUERIES\n- PromQL goes to instant_query / range_query; LogQL goes to loki_query. They are\n  different languages — do not send one to the other.\n- A LogQL query MUST carry a stream selector (e.g. '{app=\"api\"}'). A query\n  without one is rejected by the tool, not silently widened.\n- Use label_values / loki_labels / loki_label_values to discover real label\n  names and values before writing a selector. Do not guess a job, instance, or\n  app name.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, saturation, or regression unless a tool result\n  supports it.\n- Do not add generic advice that does not follow from the tool output.\n- Keep the identifiers straight: an alertname is not a label value; a silence ID\n  is not a dashboard UID; a \"job\" is a scrape config name while an \"instance\" is\n  a single scraped endpoint.\n```\n\n## Recommended setup for a local model\n\nThere is no read-only switch to set — this tool does not decide whether a\nwrite is permitted. If you want the connection to be read-only until you trust\nthe setup, enforce it at the account: give it a Grafana token with only Viewer\nscope, and a Prometheus/Alertmanager reached without the admin/write API. Any\nwrite attempt then fails at the server, which is the place that actually owns\nthe permission.\n\n```bash\nobservability-aiops doctor\n```\n\nWhen you are ready to allow writes (silences, annotations, dashboards), connect\nwith a token that has write scope, and optionally name yourself on the audit\nrow — it is an annotation, not a gate:\n\n```bash\nexport OBSERVABILITY_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport OBSERVABILITY_AUDIT_RATIONALE=\"silencing NodeDiskFilling during the 2026-07-20 disk swap\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the analysis tools —\n  `firing_alert_rca`, `target_scrape_health_analysis`,\n  `alert_noise_and_flap_analysis`, `log_error_burst_rca`, and\n  `alert_log_context` do the multi-step correlation inside one call, so the\n  model does not have to chain reads and keep alertnames, jobs, and selectors\n  straight across turns.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions, scope PromQL and LogQL with real label matchers, and use `--limit`\n  deliberately rather than pulling whole label sets or metric-name lists.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\n## Verification status\n\nThese tools have been exercised against a real stack — Prometheus 3.x,\nAlertmanager, and Grafana 13 — covering firing-alert and scrape-target RCA,\ngoverned silence and dashboard writes, and undo replay. The behaviours described\nabove are observed, not only unit-tested.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Observability-AIops](https://github.com/AIops-tools/Observability-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.2:references/capabilities.md\n\n# observability-aiops capability matrix\n\n> **39 MCP tools** (32 read, 7 write) across Prometheus\n> (HTTP API + PromQL, default port 9090, optional bearer token), a companion\n> Alertmanager (`/api/v2`, port 9093), Grafana (HTTP API, port 3000, required\n> bearer token), and Grafana Loki (HTTP API, port 3100, optional bearer/basic\n> auth, optional multi-tenant `X-Scope-OrgID`). Loki is **read-only**. The\n> Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md).\n\n## Metrics — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `instant_query` | `/api/v1/query` | PromQL evaluated at one instant (samples: metric + value + timestamp) |\n| `range_query` | `/api/v1/query_range` | PromQL over a time range (per-series point arrays) |\n| `label_values` | `/api/v1/label/<name>/values` | distinct values of a label (default `__name__` = all metric names) |\n| `series_metadata` | `/api/v1/series` | series (label-set) metadata for a selector |\n\n## Targets & status — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_targets` | `/api/v1/targets` | active scrape targets (job, instance, health, lastError), optional up/down filter |\n| `target_scrape_health` | `/api/v1/targets` | up/down summary + the unhealthy targets |\n| `dropped_targets` | `/api/v1/targets` | targets discovered but dropped by relabeling |\n| `prometheus_config_status` | `/api/v1/status/config` | running-config fingerprint (sha256) + size — never the raw YAML/secrets |\n| `prometheus_tsdb_status` | `/api/v1/status/tsdb` | TSDB head cardinality + top metrics by series count |\n\n## Rules — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_rules` | `/api/v1/rules` | recording + alerting rules (name, type, expr, health), optional type filter |\n| `rule_health` | `/api/v1/rules` | rule-evaluation health summary + erroring rules |\n\n## Alerts — Prometheus + Alertmanager (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `firing_alerts` | `/api/v1/alerts` | firing Prometheus rule alerts, grouped by severity |\n| `pending_alerts` | `/api/v1/alerts` | pending (not-yet-firing) rule alerts |\n| `alertmanager_alerts` | AM `/api/v2/alerts` | alerts as Alertmanager sees them (post grouping/silence/inhibit) |\n| `list_silences` | AM `/api/v2/silences` | silences (active, pending, expired) with matchers |\n\n## Grafana (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_dashboards` | `/api/search?type=dash-db` | dashboards (uid, title, folder, tags), optional title query |\n| `get_dashboard` | `/api/dashboards/uid/{uid}` | one dashboard's summary (title, version, panel + tag counts) |\n| `list_datasources` | `/api/datasources` | datasources (id, uid, name, type, default flag) |\n| `datasource_health` | `/api/datasources/{id}/health` | one datasource's health (status, message) |\n| `list_folders` | `/api/folders` | Grafana folders |\n\n## Loki — logs (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `loki_labels` | `/loki/api/v1/labels` | distinct label names in the lookback window |\n| `loki_label_values` | `/loki/api/v1/label/<name>/values` | distinct values of one label (name percent-encoded) |\n| `loki_query` | `/loki/api/v1/query_range` | bounded LogQL passthrough — **requires a stream selector**; lookback capped at `MAX_LOOKBACK_HOURS=24`, lines clamped to `MAX_LINE_LIMIT=1000` (default 100) |\n| `loki_tail_errors` | `/loki/api/v1/query_range` | canned error-level read for a selector (line-filter on `(?i)(error\\|fatal\\|panic\\|exception\\|traceback\\|stacktrace)`) |\n\nBounding gate: a LogQL query with **no `{…}` stream selector**, an empty query,\nor a lookback beyond the cap is rejected up front with a teaching error — no\nunbounded scan is ever issued. Label values interpolated into a selector are\nbackslash-escaped; label names in a path are percent-encoded. Loki auth is\noptional (bearer or, per target, `basic` with a `user:password` secret) and a\nmulti-tenant `X-Scope-OrgID` header is sent when the target sets `org_id`.\n\n## Overview & flagship analyses (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `observability_overview` | platform-aware | Prometheus: firing count + targets up/down + rules erroring; Grafana: dashboard/datasource/folder counts; Loki: label-name count |\n| `firing_alert_rca` | firing alerts + alerting rules | each firing alert joined to its rule expr, ranked by severity, mapped to a likely **cause + action** |\n| `target_scrape_health_analysis` | active targets | down/erroring scrapes ranked, each `lastError` classified (refused/timeout/auth/DNS/TLS) with a fix |\n| `alert_noise_and_flap_analysis` | alert instances | alertnames with many instances / exact duplicates flagged with a group_by / inhibition / longer-`for` recommendation |\n\n## Loki — log analyses & cross-signal (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `log_error_burst_rca` | selector + window (pulls current + baseline error streams) | per-stream error counts vs a baseline window; each burst classified **new_signature** (baseline 0), **volume_spike** (>= `burst_ratio`×baseline), or **single_instance** (localized to one pod/instance) with a cause + action + sample lines |\n| `log_volume_analysis` | selector + window (pulls streams + `index/stats`) | top streams by line volume, high-cardinality (high-churn) label warnings, and a retention hint from total ingest bytes |\n| `alert_log_context` | firing alertname (Prometheus) + Loki target | maps the alert's labels → a Loki stream selector (intersect with namespace/job/service/app/container/pod/instance/component, first 4 in priority order; values escaped) and returns the correlated error streams. **Best-effort**: only labels the alert and Loki share will match |\n\n## Undo (read)\n\n| Tool | Inputs | Returns |\n|------|--------|---------|\n| `undo_list` | local undo store (`limit`) | recorded, not-yet-applied reversible writes: undoId, ts, originalTool, inverseTool, note |\n\n## Writes (governed)\n\n| Tool | Risk | API path | Notes |\n|------|------|----------|-------|\n| `create_silence` | **med** | AM `POST /api/v2/silences` | **time-boxed** (requires minutes > 0); returns silenceId; undo → `expire_silence` |\n| `expire_silence` | **med** | AM `DELETE /api/v2/silence/{id}` | inverse of create_silence |\n| `create_annotation` | **medium** | `POST /api/annotations` | Grafana event marker |\n| `update_dashboard` | **med** | `POST /api/dashboards/db` | GETs the prior model first → captures it for a restore undo |\n| `delete_dashboard` | **HIGH** | `DELETE /api/dashboards/uid/{uid}` | `dry_run`; captures prior model **BEFORE** delete; undo → recreate |\n| `reload_prometheus_config` | **med** | `POST /-/reload` | records the pre-reload config hash; no undo (re-apply the prior config file) |\n| `undo_apply` | **med** | dispatches the recorded inverse tool | executes a recorded inverse; the inverse runs through its own governed tool (its real risk tier applies); single-use token; supports `dry_run` |\n\n**No Loki writes.** Loki exposes no safe operational write surface here (no\nsilence/annotation analogue), so this tool ships Loki as read-only by design.\n\n## Out of scope (by design)\n\n- **Hosted/SaaS monitoring** — Datadog, New Relic, and enterprise NMS (only\n  self-hosted Prometheus + Grafana + Loki here)\n- **Loki writes / ingestion / deletes** — read-only LogQL only (no push, no\n  delete-series, no ruler/config writes)\n- **Creating/editing Prometheus rules or scrape config**, and provisioning\n  Grafana datasources/dashboards from scratch (beyond update/delete of an existing\n  dashboard)\n- **Long-term-storage query fan-out** (Thanos/Cortex/Mimir) — the single\n  Prometheus HTTP API only\n\nFile v0.10.2:references/cli-reference.md\n\n# observability-aiops CLI reference\n\n> Covers Prometheus (HTTP API + PromQL), a companion Alertmanager, Grafana\n> (HTTP API), and Grafana Loki (LogQL, read-only). The Prometheus/Alertmanager/\n> Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack;\n> the Loki surface has not (see docs/VERIFICATION.md). The CLI is a convenience\n> subset — the full 39-tool surface is via the MCP server\n> (`observability-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nobservability-aiops init                      # interactive wizard (asks for the platform: prometheus/grafana/loki)\nobservability-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   Prometheus: /api/v1/status/buildinfo · Grafana: /api/health\n                                           #   Loki: /ready + /loki/api/v1/status/buildinfo\nobservability-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.observability-aiops/secrets.enc)\n\n```bash\nobservability-aiops secret set <target> [--value <token>]   # store bearer token (hidden prompt if no --value)\nobservability-aiops secret list                             # names only — secrets never shown\nobservability-aiops secret rm <target>\nobservability-aiops secret migrate                          # import legacy plaintext env (OBSERVABILITY_<TARGET>_TOKEN)\nobservability-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nobservability-aiops overview [--target <t>]   # snapshot: firing alerts + targets up/down + rules erroring (Prometheus)\n                                           #   or dashboard/datasource/folder counts (Grafana) / label-name count (Loki)\n```\n\n## Query (Prometheus PromQL)\n\n```bash\nobservability-aiops query instant 'up'                      # PromQL instant query\nobservability-aiops query range 'rate(x[5m])' --start ... --end ... [--step 60s]\nobservability-aiops query labels [__name__]                 # distinct label values (default = all metric names)\n```\n\n## Logs (Grafana Loki, read-only, bounded)\n\n```bash\nobservability-aiops logs labels [--hours 1] [--target <t>]              # distinct Loki label names in the window\nobservability-aiops logs query '{app=\"api\"} |= \"error\"' [--hours 1] [--limit 100]   # bounded LogQL (stream selector required)\nobservability-aiops logs errors '{app=\"api\"}' [--hours 1] [--limit 100]  # canned error-level tail for a selector\n```\n\n## Alerts\n\n```bash\nobservability-aiops alert firing [--target <t>]             # firing Prometheus rule alerts, by severity\nobservability-aiops alert silences [--target <t>]           # Alertmanager silences\nobservability-aiops alert rca [--target <t>]                # root-cause firing alerts (join to rule expr → cause+action)\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `query`, `logs`, and `alert` are the CLI subset; the remaining\n  metrics, targets, rules, Grafana, Loki analyses (log_error_burst_rca,\n  log_volume_analysis, alert_log_context), and governed-write tools\n  (create/expire silence, create annotation, update/delete dashboard, reload\n  config) are exposed through the MCP server. High-risk MCP writes honour `OBSERVABILITY_AUDIT_APPROVED_BY` /\n  `OBSERVABILITY_AUDIT_RATIONALE` and support dry-run.\n\nFile v0.10.2:references/setup-guide.md\n\n# observability-aiops setup & security guide\n\n> The Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md). **Prometheus,\n> Grafana, and Loki are all free/open-source and trivial to stand up in a lab\n> (`docker run prom/prometheus`, `grafana/grafana`, `grafana/loki`), so a live\n> `doctor` check is easy.**\n\n## 1. Install\n\n```bash\nuv tool install observability-aiops\n```\n\n## 2. Get a credential\n\n- **Prometheus** — a bearer token is **optional**; many self-hosted deployments\n  are unauthenticated. observability-aiops talks to the HTTP API on port **9090**\n  and a companion Alertmanager on **9093** (`/api/v2`).\n- **Grafana** — a **service-account token** (Administration → Service accounts →\n  Add token) or legacy API key is **required**. Grafana's HTTP API is on port\n  **3000**.\n- **Loki** — auth is **optional** (bearer token, or per-target `basic` auth where\n  the stored secret is `user:password`). The HTTP API is on port **3100**. For a\n  multi-tenant deployment set `org_id` to send the `X-Scope-OrgID` header. Loki is\n  **read-only** here.\n\n## 3. Onboard\n\n```bash\nobservability-aiops init\n```\n\nThe wizard asks, per target, for the **platform** (`prometheus` / `grafana` /\n`loki`), the **host**, the **scheme** (`http` / `https`), the **port** (defaults\n9090 for Prometheus, 3000 for Grafana, 3100 for Loki), an optional **Alertmanager\nURL** (Prometheus only), the **auth type** and **org id** (Loki only), and the\n**token** — required for Grafana, optional for Prometheus/Loki. Non-secret\nconnection details go to `~/.observability-aiops/config.yaml`; the token is stored\n**encrypted** into `~/.observability-aiops/secrets.enc`. Example config (one config\ncan span the whole stack):\n\n```yaml\ntargets:\n  - name: prod-prom\n    platform: prometheus\n    host: 10.0.0.20\n    scheme: http\n    port: 9090\n    alertmanager_url: http://10.0.0.20:9093   # optional; blank assumes host:9093\n  - name: prod-grafana\n    platform: grafana\n    host: 10.0.0.30\n    scheme: https\n    port: 3000\n    verify_ssl: true\n  - name: prod-loki\n    platform: loki\n    host: 10.0.0.40\n    scheme: http\n    port: 3100\n    auth_type: bearer        # or 'basic' (secret is user:password)\n    org_id: team-a           # optional; sent as X-Scope-OrgID (multi-tenant)\n```\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nExport the master password so the encrypted store can be unlocked without a\nprompt:\n\n```bash\nexport OBSERVABILITY_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## Credential security\n\n- The token is **never** written to disk in plaintext. It lives only in\n  `~/.observability-aiops/secrets.enc`, encrypted with Fernet (AES-128-CBC +\n  HMAC), the key derived from your master password via scrypt. Only a per-store\n  random salt and the ciphertext are on disk (chmod 600); the master password\n  itself is never stored.\n- A legacy plaintext env var `OBSERVABILITY_<TARGET_NAME_UPPER>_TOKEN` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `observability-aiops secret migrate` (it imports then renames the old `.env`).\n- The token is sent as an `Authorization: Bearer` header at request time and held\n  only in memory; it is never logged or echoed. Exception text and tracebacks are\n  scrubbed of secret-shaped strings before being written to the audit log.\n\n## Governance harness state\n\nState lives under `~/.observability-aiops/` (relocate with `OBSERVABILITY_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with risk tier and any operator-supplied\n  approver/rationale (optional annotations, never required)\n- `undo.db` — inverse descriptors for reversible writes (create_silence→expire,\n  update/delete dashboard→restore/recreate)\n- budget / runaway guard — caps cumulative tool calls and wall-time; trips on\n  tight poll/retry loops\n\n## Governed writes\n\n- **High-risk** op (`delete_dashboard`) supports `dry_run` and captures the full\n  prior dashboard model **before** deleting so the recorded undo can recreate\n  it. Optionally set `OBSERVABILITY_AUDIT_APPROVED_BY` and\n  `OBSERVABILITY_AUDIT_RATIONALE` to annotate the audit row — neither is\n  required, and the write runs either way.\n- **Reversible** writes capture the real fetched before-state:\n  `update_dashboard` (restore prior model), `create_silence` (expire the created\n  silence). `reload_prometheus_config` records the pre-reload config hash.\n- **Time-boxed** ops require a positive duration: `create_silence` (in minutes).\n  This prevents forgotten, indefinite silences.\n\n## Verify\n\n```bash\nobservability-aiops doctor\n```\n\n`doctor` is platform-aware: it checks the config file, the encrypted store and its\npermissions, that a token is present where required, and (unless `--skip-auth`)\nconnectivity — `/api/v1/status/buildinfo` for Prometheus targets, `/api/health`\nfor Grafana targets, and `/ready` + `/loki/api/v1/status/buildinfo` for Loki\ntargets.\n\n## Loki query bounding (safety)\n\nLoki reads are deliberately bounded so an agent can't ask for \"all logs, forever\":\n\n- Every `loki_query` must carry a `{…}` **stream selector** — an unbounded query\n  with no selector (or an empty query) is rejected with a teaching error.\n- Lookback is capped at **24h** (`MAX_LOOKBACK_HOURS`); a longer window is refused.\n- Returned lines are clamped to **1000** (`MAX_LINE_LIMIT`, default 100).\n- `loki_tail_errors` wraps a selector with a canned case-insensitive error filter.\n- Label values interpolated into a selector are backslash-escaped and label names\n  in a path segment are percent-encoded, so a hostile label value can't break out\n  of the LogQL string or rewrite the request path.\n\nFile v0.10.2:skill-card.md\n\n## Description:\n\nObservability AIops helps age\n\nArchive v0.10.1: 7 files, 21370 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (3052b), SKILL.md (20275b), _meta.json (139b)\n\nArchive v0.10.0: 7 files, 21118 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2833b), SKILL.md (19949b), _meta.json (139b)\n\nArchive v0.9.0: 7 files, 21174 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2960b), SKILL.md (20075b), _meta.json (138b)\n\nArchive v0.8.0: 7 files, 21078 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2784b), SKILL.md (20075b), _meta.json (138b)\n\nArchive v0.7.0: 7 files, 21109 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2860b), SKILL.md (20075b), _meta.json (138b)\n\nArchive v0.6.0: 7 files, 21121 bytes\n\nFiles: references/agent-guardrails.md (7225b), references/capabilities.md (7902b), references/cli-reference.md (3495b), references/setup-guide.md (5752b), skill-card.md (2760b), SKILL.md (20075b), _meta.json (138b)\n\nArchive v0.5.0: 7 files, 20884 bytes\n\nFiles: references/agent-guardrails.md (7022b), references/capabilities.md (7927b), references/cli-reference.md (3495b), references/setup-guide.md (5700b), skill-card.md (2994b), SKILL.md (19606b), _meta.json (138b)","readmeExcerpt":"Skill: observability-aiops Owner: zw008 Summary: Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager al","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install observability-aiops\nobservability-aiops init       # wizard: pick platform (prometheus/grafana) + encrypted token\nobservability-aiops doctor"},{"language":"bash","snippet":"openclaw plugins install clawhub:@zw008/observability-aiops\nopenclaw skills info observability-aiops          # expect: Visible to model: yes"},{"language":"text","snippet":"You operate a self-hosted observability stack (Prometheus, Alertmanager,\nGrafana, Loki) through the observability-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current state of the stack, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit (or a narrower\n  selector) instead of treating the partial result as complete.\n- A null field means the API did not return that value. Report it as \"not\n  available\" — never infer it. A missing \"severity\" label is not \"info\".\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  alert names, severities, label values, or health strings.\n- When an analysis tool returns ranked findings, work in the order given and\n  cite the measured number the ranking is based on.\n\nQUERIES\n- PromQL goes to instant_query / range_query; LogQL goes to loki_query. They are\n  different languages — do not send one to the other.\n- A LogQL query MUST carry a stream selector (e.g. '{app=\"api\"}'). A query\n  without one is rejected by the tool, not silently widened.\n- Use label_values / loki_labels / loki_label_values to discover real label\n  names and values before writing a selector. Do not guess a job, instance, or\n  app name.\n\n- Apart from `undo apply`, no write tool here has a CLI command, so nothing will ask you\n  to confirm one —\n  `delete_dashboard` included. Call with `dry_run=True` first, show the operator what would\n  change, and wait for an explicit go-ahead.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do"},{"language":"bash","snippet":"observability-aiops doctor"},{"language":"bash","snippet":"export OBSERVABILITY_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport OBSERVABILITY_AUDIT_RATIONALE=\"silencing NodeDiskFilling during the 2026-07-20 disk swap\""},{"language":"bash","snippet":"observability-aiops init                      # interactive wizard (asks for the platform: prometheus/grafana/loki)\nobservability-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   Prometheus: /api/v1/status/buildinfo · Grafana: /api/health\n                                           #   Loki: /ready + /loki/api/v1/status/buildinfo\nobservability-aiops mcp                       # start the MCP server (stdio transport)"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: observability-aiops\nslug: observability-aiops\ndisplayName: \"Observability AIops\"\nsummary: \"Governed Prometheus + Grafana ops: PromQL, alerts, dashboards, RCA; 39 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Observability-AIops\ntags: [aiops, mcp, governance, observability]\ndescription: >\n  Use this skill whenever the user needs to operate a self-hosted observability stack on Prometheus (HTTP API + PromQL), Alertmanager, Grafana, or Grafana Loki (logs) — a one-shot overview, PromQL instant/range queries, label + series metadata, scrape-target health (up/down + why) and dropped targets, recording/alerting rule health, firing/pending alerts, Alertmanager alerts + silences, Grafana dashboards/datasources/folders, bounded Loki LogQL log reads (labels, query, error-tail), five flagship analyses (firing-alert RCA, target-scrape-health, alert-noise/flap, log-error-burst RCA, log-volume/cardinality) plus an alert->log cross-signal, and guarded writes (create/expire silence, create annotation, update/delete dashboard, reload Prometheus config).\n  Always use this skill for \"Prometheus\", \"PromQL\", \"Alertmanager\", \"Grafana\", \"Loki\", \"LogQL\", \"logs\", \"which targets are down\", \"scrape failing\", \"why is this alert firing\", \"root cause this alert\", \"firing alerts\", \"silence this alert\", \"noisy alerts\", \"alert flapping\", \"recording rule\", \"alerting rule\", \"dashboard\", \"datasource health\", \"reload prometheus config\", \"TSDB cardinality\", \"error burst\", \"log volume\", \"log cardinality\", \"tail errors\" when the context is a self-hosted metrics/logs/observability stack.\n  Do NOT use when the target is something other than a Prometheus/Grafana observability stack (a hypervisor, storage appliance, backup product, container-orchestrator control plane, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.\n  Governed observability operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack (RCAs, governed writes, undo); the Loki surface has not (see docs/VERIFICATION.md).\ninstaller:\n  kind: uv\n  package: observability-aiops\nargument-hint: \"[a PromQL query, an alert/dashboard uid, or describe your observability task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"observability-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"OBSERVABILITY_AIOPS_CONFIG\",\"OBSERVABILITY_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Observability-AIops\",\"emoji\":\"📈\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed observability operations across Prometheus (HTTP API + PromQL, default port 9090, optional bearer token), a companion Alertmanager (/api/v2, default port 9093), Grafana (HTTP API, default"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"observability-aiops\",\n  \"version\": \"0.10.4\",\n  \"publishedAt\": 1789601183574\n}"},{"path":"references/agent-guardrails.md","content":"# Agent guardrails — running observability-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## What the tool now enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Don't invent a value when a field is missing\" | A field the API did not return comes back as `null`, never as `\"\"`. An alert with no `severity` label, a scrape target that has never errored, a recording rule with no alert `state`, a silence with no `comment` — all report `null`, distinguishable from a genuinely empty value. |\n| \"Tell me if the output was cut off\" | Bounded reads (`loki_query`, `loki_tail_errors`, `loki_labels`, `loki_label_values`, `instant_query`, `range_query`, `label_values`, `series_metadata`, `undo_list`) return `{\"returned\": N, \"limit\": L, \"truncated\": true/false}` alongside the rows. For the Loki reads truncation is **measured** — one line beyond the limit is requested — not guessed from a length coincidence. |\n| \"Preserve the ordering / tell me what's most urgent\" | The analysis tools already return worst-first: `firing_alert_rca` ranks by severity, `target_scrape_health_analysis` puts down targets before slow ones, `alert_noise_and_flap_analysis` sorts by instance count. Priority is the list order, and each entry carries the measured number it was ranked on. |\n| \"Confirm before anything destructive\" | Every write tool takes `dry_run` for a preview, and `delete_dashboard` is `risk=high`. ⚠️ **Apart from `undo apply`, no write tool here has a CLI command**, so the CLI double confirmation never applies to one: they are reachable only over MCP, where nothing prompts. Keep your own confirmation for them. |\n| \"Log what you did\" | Every call is audited to `~/.observability-aiops/audit.db` regardless of what the model says it did, and reversible writes record an undo token (`undo_list` / `undo_apply`). |\n| \"Don't hammer the same call in a loop\" | The runaway guard trips a circuit breaker on tight poll/retry loops — a safety backstop, not authorization. |\n\nAuthorization is not this tool's job. Whether a write is allowed to happen is\ndecided by the account you connect it with, or by your agent's own judgement —\nnot by this harness. See \"Recommended setup for a local model\" below for how to\nenforce read-only at the account instead of in a prompt.\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYo"},{"path":"references/capabilities.md","content":"# observability-aiops capability matrix\n\n> **39 MCP tools** (32 read, 7 write) across Prometheus\n> (HTTP API + PromQL, default port 9090, optional bearer token), a companion\n> Alertmanager (`/api/v2`, port 9093), Grafana (HTTP API, port 3000, required\n> bearer token), and Grafana Loki (HTTP API, port 3100, optional bearer/basic\n> auth, optional multi-tenant `X-Scope-OrgID`). Loki is **read-only**. The\n> Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not\n> (see docs/VERIFICATION.md).\n\n## Metrics — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `instant_query` | `/api/v1/query` | PromQL evaluated at one instant (samples: metric + value + timestamp) |\n| `range_query` | `/api/v1/query_range` | PromQL over a time range (per-series point arrays) |\n| `label_values` | `/api/v1/label/<name>/values` | distinct values of a label (default `__name__` = all metric names) |\n| `series_metadata` | `/api/v1/series` | series (label-set) metadata for a selector |\n\n## Targets & status — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_targets` | `/api/v1/targets` | active scrape targets (job, instance, health, lastError), optional up/down filter |\n| `target_scrape_health` | `/api/v1/targets` | up/down summary + the unhealthy targets |\n| `dropped_targets` | `/api/v1/targets` | targets discovered but dropped by relabeling |\n| `prometheus_config_status` | `/api/v1/status/config` | running-config fingerprint (sha256) + size — never the raw YAML/secrets |\n| `prometheus_tsdb_status` | `/api/v1/status/tsdb` | TSDB head cardinality + top metrics by series count |\n\n## Rules — Prometheus (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_rules` | `/api/v1/rules` | recording + alerting rules (name, type, expr, health), optional type filter |\n| `rule_health` | `/api/v1/rules` | rule-evaluation health summary + erroring rules |\n\n## Alerts — Prometheus + Alertmanager (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `firing_alerts` | `/api/v1/alerts` | firing Prometheus rule alerts, grouped by severity |\n| `pending_alerts` | `/api/v1/alerts` | pending (not-yet-firing) rule alerts |\n| `alertmanager_alerts` | AM `/api/v2/alerts` | alerts as Alertmanager sees them (post grouping/silence/inhibit) |\n| `list_silences` | AM `/api/v2/silences` | silences (active, pending, expired) with matchers |\n\n## Grafana (read)\n\n| Tool | API path | Returns |\n|------|----------|---------|\n| `list_dashboards` | `/api/search?type=dash-db` | dashboards (uid, title, folder, tags), optional title query |\n| `get_dashboard` | `/api/dashboards/uid/{uid}` | one dashboard's summary (title, version, panel + tag counts) |\n| `list_datasources` | `/api/datasources` | datasources (id, uid, name, type, default flag) |\n| `datasource_health` | `/api/datasources/{id}/health` | one datasource's health (status"},{"path":"references/cli-reference.md","content":"# observability-aiops CLI reference\n\n> Covers Prometheus (HTTP API + PromQL), a companion Alertmanager, Grafana\n> (HTTP API), and Grafana Loki (LogQL, read-only). The Prometheus/Alertmanager/\n> Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack;\n> the Loki surface has not (see docs/VERIFICATION.md). The CLI is a convenience\n> subset — the full 39-tool surface is via the MCP server\n> (`observability-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nobservability-aiops init                      # interactive wizard (asks for the platform: prometheus/grafana/loki)\nobservability-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   Prometheus: /api/v1/status/buildinfo · Grafana: /api/health\n                                           #   Loki: /ready + /loki/api/v1/status/buildinfo\nobservability-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.observability-aiops/secrets.enc)\n\n```bash\nobservability-aiops secret set <target> [--value <token>]   # store bearer token (hidden prompt if no --value)\nobservability-aiops secret list                             # names only — secrets never shown\nobservability-aiops secret rm <target>\nobservability-aiops secret migrate                          # import legacy plaintext env (OBSERVABILITY_<TARGET>_TOKEN)\nobservability-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nobservability-aiops overview [--target <t>]   # snapshot: firing alerts + targets up/down + rules erroring (Prometheus)\n                                           #   or dashboard/datasource/folder counts (Grafana) / label-name count (Loki)\n```\n\n## Query (Prometheus PromQL)\n\n```bash\nobservability-aiops query instant 'up'                      # PromQL instant query\nobservability-aiops query range 'rate(x[5m])' --start ... --end ... [--step 60s]\nobservability-aiops query labels [__name__]                 # distinct label values (default = all metric names)\n```\n\n## Logs (Grafana Loki, read-only, bounded)\n\n```bash\nobservability-aiops logs labels [--hours 1] [--target <t>]              # distinct Loki label names in the window\nobservability-aiops logs query '{app=\"api\"} |= \"error\"' [--hours 1] [--limit 100]   # bounded LogQL (stream selector required)\nobservability-aiops logs errors '{app=\"api\"}' [--hours 1] [--limit 100]  # canned error-level tail for a selector\n```\n\n## Alerts\n\n```bash\nobservability-aiops alert firing [--target <t>]             # firing Prometheus rule alerts, by severity\nobservability-aiops alert silences [--target <t>]           # Alertmanager silences\nobservability-aiops alert rca [--target <t>]                # root-cause firing alerts (join to rule expr → cause+action)\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target d"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1927,"uniquenessScore":40,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T11:32:30.020Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T13:32:00.088Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}