{"id":"3139bd5b-58c8-4789-80ec-a3d703c05cfb","entityType":"agent","slug":"clawhub-zw008-monitoring-aiops","name":"monitoring-aiops","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-monitoring-aiops","canonicalPath":"/agent/clawhub-zw008-monitoring-aiops","generatedAt":"2026-10-10T09:09:08.206Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":null},"description":"Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window). Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring. Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.7K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:monitoring-aiops","sourceUrl":"https://clawhub.ai/zw008/monitoring-aiops","homepage":"https://clawhub.ai/zw008/skills/monitoring-aiops","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/monitoring-aiops","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/monitoring-aiops","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":65,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"monitoring-aiops technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":null},"stars":null,"forks":null,"downloads":1691,"packageName":null,"latestVersion":"0.10.4","tractionLabel":"1.7K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T04:22:24.953Z","lastCrawledAt":"2026-10-10T04:22:24.953Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T04:22:24.953Z","lastVerifiedAt":null,"highlights":[{"version":"0.10.4","createdAt":"2026-09-16T23:25:58.731Z","changelog":"monitoring-aiops 0.10.4 - Updated agent guardrails documentation in references/agent-guardrails.md. - Removed redundant skill-card.md file. - No changes to functional behavior or features.","fileCount":7,"zipByteSize":19525},{"version":"0.10.3","createdAt":"2026-09-15T06:06:01.725Z","changelog":"- Removed the skill-card.md file from the project. - No new features or bug fixes; this is a documentation cleanup release. - All functionality and interfaces remain unchanged.","fileCount":7,"zipByteSize":19423},{"version":"0.10.2","createdAt":"2026-09-12T14:29:43.973Z","changelog":"monitoring-aiops 0.10.2 - Documentation cleanup: removed redundant file (`skill-card.md`) to reduce duplication. - Kept SKILL.md as the single authoritative documentation source. - No changes to core features, operations, or supported platforms.","fileCount":7,"zipByteSize":19394},{"version":"0.10.1","createdAt":"2026-09-12T10:13:35.678Z","changelog":"- Removed legacy file skill-card.md from the project. - Updated documentation in SKILL.md with no functional changes to the skill. - No new features or breaking changes introduced in this version.","fileCount":7,"zipByteSize":19518},{"version":"0.10.0","createdAt":"2026-09-12T01:02:15.524Z","changelog":"monitoring-aiops v0.10.0 - Updated metadata in SKILL.md: improved environment variable and binary specification for flexibility (`anyBins`, optional envs). - Minor changes to internal metadata formatting for better compatibility. - Removed `skill-card.md` file. - No changes to core functionality or tool coverage.","fileCount":7,"zipByteSize":19222},{"version":"0.9.0","createdAt":"2026-08-10T06:52:14.099Z","changelog":"- Removed the file: skill-card.md. - No functional or user-facing changes; only documentation cleanup.","fileCount":7,"zipByteSize":19242},{"version":"0.8.0","createdAt":"2026-08-03T05:53:32.911Z","changelog":"- Removed the sample file skill-card.md. - No functional or operational changes to skill logic. - All documented capabilities and platform compatibility remain unchanged.","fileCount":7,"zipByteSize":19147},{"version":"0.7.0","createdAt":"2026-08-02T09:40:17.412Z","changelog":"- Removed the file: skill-card.md - No changes to core functionality—this is a documentation cleanup release. - All features, compatibility, and governance details remain unchanged.","fileCount":7,"zipByteSize":19167}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:monitoring-aiops","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:monitoring-aiops` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/monitoring-aiops before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T09:09:08.202Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-monitoring-aiops/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":null},"readme":"Skill: monitoring-aiops\n\nOwner: zw008\n\nSummary: Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window). Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring. Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill. Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.\n\nTags: agent-skills:0.1.0, ai-ops:0.1.0, latest:0.10.4, mcp:0.1.0, prtg:0.1.0, solarwinds:0.1.0\n\nVersion history:\n\nv0.10.4 | 2026-09-16T23:25:58.731Z | auto\n\nmonitoring-aiops 0.10.4\n\n- Updated agent guardrails documentation in references/agent-guardrails.md.\n- Removed redundant skill-card.md file.\n- No changes to functional behavior or features.\n\nv0.10.3 | 2026-09-15T06:06:01.725Z | auto\n\n- Removed the skill-card.md file from the project.\n- No new features or bug fixes; this is a documentation cleanup release.\n- All functionality and interfaces remain unchanged.\n\nv0.10.2 | 2026-09-12T14:29:43.973Z | auto\n\nmonitoring-aiops 0.10.2\n\n- Documentation cleanup: removed redundant file (`skill-card.md`) to reduce duplication.\n- Kept SKILL.md as the single authoritative documentation source.\n- No changes to core features, operations, or supported platforms.\n\nv0.10.1 | 2026-09-12T10:13:35.678Z | auto\n\n- Removed legacy file skill-card.md from the project.\n- Updated documentation in SKILL.md with no functional changes to the skill.\n- No new features or breaking changes introduced in this version.\n\nv0.10.0 | 2026-09-12T01:02:15.524Z | auto\n\nmonitoring-aiops v0.10.0\n\n- Updated metadata in SKILL.md: improved environment variable and binary specification for flexibility (`anyBins`, optional envs).\n- Minor changes to internal metadata formatting for better compatibility.\n- Removed `skill-card.md` file.\n- No changes to core functionality or tool coverage.\n\nv0.9.0 | 2026-08-10T06:52:14.099Z | auto\n\n- Removed the file: skill-card.md.\n- No functional or user-facing changes; only documentation cleanup.\n\nv0.8.0 | 2026-08-03T05:53:32.911Z | auto\n\n- Removed the sample file skill-card.md.\n- No functional or operational changes to skill logic.\n- All documented capabilities and platform compatibility remain unchanged.\n\nv0.7.0 | 2026-08-02T09:40:17.412Z | auto\n\n- Removed the file: skill-card.md\n- No changes to core functionality—this is a documentation cleanup release.\n- All features, compatibility, and governance details remain unchanged.\n\nv0.6.0 | 2026-07-21T09:41:46.900Z | auto\n\nmonitoring-aiops 0.6.0\n\n- Audit/governance harness policy engine simplified: risk tiers are now recorded as descriptive labels on the audit trail instead of being enforced as runtime gates.\n- Documentation updated to clarify new risk-tier labelling approach and adjusted governance harness description.\n- Removed legacy skill-card.md file.\n- Minor compatibility and metadata clarifications.\n- References and setup guides updated for consistency with new governance flow.\n\nv0.5.0 | 2026-07-20T11:16:02.106Z | auto\n\nmonitoring-aiops 0.5.0\n\n- Compatibility updated: Now supports SolarWinds Orion API port 17774 for Orion 2023.1+, with automatic fallback to legacy port 17778.\n- Documentation improved across SKILL.md, setup-guide, and capabilities references.\n- Removed deprecated skill-card.md file.\n- No user-facing functional changes to monitoring features.\n\nv0.4.2 | 2026-07-20T04:07:30.731Z | auto\n\nmonitoring-aiops v0.4.2\n\n- Removed the file: skill-card.md.\n- No changes to functionality; documentation file cleanup only.\n\nv0.4.1 | 2026-07-20T03:07:45.637Z | auto\n\n## monitoring-aiops 0.4.1\n\n- Removed the `skill-card.md` file from the project.\n- No feature or functional changes—maintenance/cleanup release only.\n\nv0.4.0 | 2026-07-19T03:52:04.223Z | auto\n\nmonitoring-aiops 0.4.0\n\n- Added agent guardrails documentation (references/agent-guardrails.md).\n- Expanded toolset: now provides 42 governed monitoring tools.\n- SKILL.md updated for improved summary, tagging, and tool count clarity.\n- Validation status and live-check info clarified.\n- Removed obsolete skill-card.md and refined documentation structure for easier use.\n\nv0.3.0 | 2026-07-17T05:55:44.323Z | auto\n\n**monitoring-aiops v0.3.0 – adds Zabbix NOC monitoring and governance**\n\n- Added Zabbix 6.x/7.x support: JSON-RPC API integration for problems, hosts, host groups, triggers, events, item history, and maintenances (create/delete).\n- Skill now spans SolarWinds Orion, Paessler PRTG, and Zabbix platforms in a unified audit-governed NOC toolset.\n- Updated all documentation and guides to include Zabbix coverage and changes to compatible platforms.\n- Expanded tool count from 31 to 40, including new Zabbix read/write operations and governance (undo, state capture, dry-run/confirmation).\n- Zabbix item history queries are time-bounded, and all maintenance actions are audited with previous state where applicable.\n- Removed the deprecated skill-card.md (documentation cleanup).\n\nv0.2.0 | 2026-07-13T13:08:50.474Z | auto\n\nmonitoring-aiops 0.2.0\n\n- Removed the skill-card.md file for cleaner packaging.\n- Updated SKILL.md for minor improvements and alignment with the latest version.\n- No changes to core functionality; documentation and packaging cleanup only.\n\nv0.1.0 | 2026-07-12T07:00:41.273Z | auto\n\nInitial preview release of monitoring-aiops, enabling governed operations across SolarWinds Orion and Paessler PRTG monitoring platforms.\n\n- Provides 31 tools for NOC monitoring and management, covering SWQL, alerts, health, and maintenance actions.\n- Supports both SolarWinds (SWIS REST + SWQL) and PRTG (web API), allowing one config to span multiple NOCs.\n- All write operations are audited and governed by an integrated policy and undo harness, with role-based risk tiers and token budgets.\n- Secrets (passwords/API tokens) are encrypted on disk and never stored in plaintext; CLI onboarding and migration included.\n- Read-only SWQL passthrough restricted to SELECT; write actions require confirmation, audit, and are dry-run validated.\n- Standalone, with no external dependencies or webhooks; preview mode with mock-validation and support for live test against PRTG Freeware.\n\nArchive index:\n\nArchive v0.10.4: 7 files, 19525 bytes\n\nFiles: references/agent-guardrails.md (7748b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2711b), SKILL.md (16987b), _meta.json (136b)\n\nFile v0.10.4:SKILL.md\n\n---\nname: monitoring-aiops\nslug: monitoring-aiops\ndisplayName: \"Monitoring AIops\"\nsummary: \"Governed SolarWinds Orion + PRTG + Zabbix ops: SWQL, alert rollup, health, 42 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Monitoring-AIops\ntags: [aiops, mcp, governance, monitoring]\ndescription: >\n  Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window).\n  Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring.\n  Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill.\n  Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.\ninstaller:\n  kind: uv\n  package: monitoring-aiops\nargument-hint: \"[node/sensor id, a SWQL question, or describe your NOC task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"monitoring-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"MONITORING_AIOPS_CONFIG\",\"MONITORING_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Monitoring-AIops\",\"emoji\":\"📡\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed monitoring operations across SolarWinds Orion (SWIS REST + SWQL, port 17774 on Orion 2023.1+ with an automatic one-shot fallback to the legacy 17778, HTTP Basic auth), Paessler PRTG (web API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at /api_jsonrpc.php, API token as Bearer header on 6.4+/7.x with a legacy auth-field fallback for 6.0). Each target in the config names its own platform, so one config can span all NOCs. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.monitoring-aiops/ (relocatable via MONITORING_AIOPS_HOME).\n  Credentials: the Orion account password (SolarWinds), the PRTG API token, or the Zabbix API token is stored ENCRYPTED in ~/.monitoring-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'monitoring-aiops init' to onboard (it asks for the platform), or 'monitoring-aiops secret set <target>' to add one. The store is unlocked by a master password from MONITORING_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var MONITORING_<TARGET_NAME_UPPER>_SECRET is still honoured as a fallback with a deprecation warning (migrate with 'monitoring-aiops secret migrate'). The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG/Zabbix API token at request time and held only in memory; secrets are never logged or echoed.\n  Read-only SWQL passthrough (swql_query) is validated to accept SELECT statements only. State-changing operations pass through the @governed_tool decorator (budget guard + audit + undo recording; each tool's risk_level is recorded as a descriptive tier, not a gate). Destructive writes (unmanage_node, remove_node, zabbix_delete_maintenance) are high-risk with dry_run + double confirmation; unmanage_node records an inverse remanage undo descriptor, zabbix_delete_maintenance captures the window's FULL definition into priorState first. Suppression/maintenance writes are TIME-BOXED (mute_alerts, schedule_maintenance, schedule_maintenance_prtg, zabbix_create_maintenance require an end time / duration). mute_alerts→unmute, pause_sensor→resume, zabbix_create_maintenance→delete-that-maintenance-id record inverse undo descriptors. Zabbix item history is BOUNDED (capped window + point count).\n  Webhooks: none — no outbound network calls beyond the configured SolarWinds SWIS / PRTG web API / Zabbix JSON-RPC endpoint.\n  SSL: verify_ssl defaults to false-friendly for self-signed lab certs; enable for production.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked SWIS/PRTG/Zabbix responses; not yet run against a live NOC (see docs/VERIFICATION.md). PRTG has a free perpetual 100-sensor Freeware edition with the API, and Zabbix is fully open source (a Docker-compose appliance is a 10-minute live check) — the easiest live checks; SolarWinds is a 30-day trial (mock-only past that — largest verification debt).\n---\n\n# Monitoring AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by SolarWinds, Paessler, Zabbix, or any monitoring vendor.** SolarWinds, Orion, SWQL, THWACK, PRTG, Paessler and Zabbix are trademarks of their respective owners. Source at [github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops) under the MIT license.\n\nGoverned network / infrastructure monitoring operations — **42 MCP tools**\nacross **SolarWinds Orion** (SWIS REST + SWQL), **Paessler PRTG** (web API),\nand **Zabbix 6.x/7.x** (JSON-RPC 2.0),\nevery one wrapped with the bundled `@governed_tool` harness: a local unified\naudit log under `~/.monitoring-aiops/`, policy engine, token/runaway budget\nguard, undo-token recording, and risk-tier labelling on the audit trail. One\nconfig can span all NOCs. The Orion password / PRTG API token / Zabbix API token is stored\n**encrypted** (`~/.monitoring-aiops/secrets.enc`, Fernet + scrypt) — never\nplaintext on disk.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`monitoring_aiops.governance`) — no external skill-family dependency.\n> PRTG's free Freeware edition and an open-source Zabbix appliance are the\n> easiest live checks; SolarWinds is trial-only past 30 days (largest\n> verification debt — see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **SWQL** | SolarWinds | library, canned, query (SELECT-only passthrough) | 3 | read |\n| **Alerts** | all | active_alerts (dedup/rollup), alert_acknowledge | 2 | 1 read, 1 write |\n| **SolarWinds health** | SolarWinds | node/nodes/interface/volume/application status, topn, noc_rollup | 7 | read |\n| **SolarWinds writes** | SolarWinds | list_events/unmanaged/muted | 3 | read |\n| | SolarWinds | mute/unmute, schedule_maintenance, remanage_node | 4 | write (med) |\n| | SolarWinds | unmanage_node, remove_node | 2 | write (**high**) |\n| **PRTG** | PRTG | sensors/sensor_details/devices/groups/history/system_status/alarms | 7 | read |\n| **PRTG writes** | PRTG | pause_sensor, resume_sensor, schedule_maintenance_prtg | 3 | write (med) |\n| **Zabbix** | Zabbix | zabbix_problems/hosts/hostgroups/triggers/events/item_history/maintenances | 7 | read |\n| **Zabbix writes** | Zabbix | zabbix_create_maintenance (time-boxed; undo = delete that id) | 1 | write (med) |\n| | Zabbix | zabbix_delete_maintenance (priorState = full definition) | 1 | write (**high**) |\n| **Undo** | all | undo_list, undo_apply | 2 | undo |\n\nThe canned SWQL library (`swql_library` lists them) answers the most-repeated\nTHWACK questions directly: `nodes_down`, `flapping_interfaces`, `muted_report`,\n`high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled`. For anything else,\n`swql_query` is a validated read-only (SELECT-only) SWQL passthrough.\n\n## Quick Install\n\n```bash\nuv tool install monitoring-aiops\nmonitoring-aiops init       # wizard: pick platform (solarwinds/prtg/zabbix) + encrypted secret\nmonitoring-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/monitoring-aiops\nopenclaw skills info monitoring-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a NOC snapshot (`overview` / `noc_rollup`): active/unacked alert counts,\n  down/warning nodes, worst CPU\n- Answer a repeated SWQL question (`swql_library` → `swql_canned nodes_down`), or\n  run an ad-hoc read-only SWQL SELECT (`swql_query`)\n- Triage an alert storm (`active_alerts` dedup/rollup collapses flap/down\n  storms), then `alert_acknowledge`\n- SolarWinds health: `node_status`, `interface_status` (top-N by util),\n  `volume_status`, `application_status` (SAM), `topn` (cpu/mem/latency/loss)\n- PRTG: list `prtg_sensors` / `prtg_devices` / `prtg_groups`, drill with\n  `prtg_sensor_details` / `prtg_history`, check `prtg_alarms` / `prtg_system_status`\n- Zabbix: triage `zabbix_problems` (0-5 severity mapped to levels) /\n  `zabbix_triggers`, inventory `zabbix_hosts` / `zabbix_hostgroups`, drill with\n  `zabbix_item_history` (bounded), review `zabbix_events` / `zabbix_maintenances`\n- Safely take a node out for maintenance (`schedule_maintenance` /\n  `unmanage_node` with dry_run + double-confirm), pause a PRTG sensor\n  (`pause_sensor`), or create a time-boxed Zabbix maintenance window\n  (`zabbix_create_maintenance` — undo deletes exactly that window)\n\n**Do NOT use when** the target is not a SolarWinds/PRTG/Zabbix monitoring\nplatform — route hypervisor, storage, backup, cluster, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| SolarWinds Orion / SWQL, PRTG, or Zabbix monitoring ops | **monitoring-aiops** (this skill) |\n| A non-monitoring platform (hypervisor, storage, backup, cluster, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Other monitoring stacks (not SolarWinds/PRTG/Zabbix) | out of scope for this tool |\n\n## Common Workflows\n\n> **No authorization gate**: the skill runs the operations you ask for and audits every one; it does not decide whether a write is permitted — that is the agent's judgement or the permissions of the SolarWinds/PRTG/Zabbix account it connects with (a read-only monitoring account makes writes fail at the server). There is no read-only switch, policy file, or approval gate. `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n\n### 1. The 3 a.m. alert storm — collapse it, then acknowledge what matters\n\n1. `monitoring-aiops doctor` → confirm the NOC platform is actually reachable\n   (a \"storm\" is sometimes just a poller that lost the target)\n2. `monitoring-aiops overview` → the one-screen picture: down/warning counts\n   across the configured targets\n3. `monitoring-aiops alert list` (MCP: `active_alerts`) → deduped / rolled-up\n   entries; an interface-flap or node-down storm collapses into **one** entry\n   with a count instead of a wall of alerts\n4. `noc_rollup` → confirm whether the storm has a single upstream cause (one\n   node down taking its children with it) rather than N independent faults\n5. Acknowledge only the rolled-up entry that matters:\n   `monitoring-aiops alert ack <alert-id>` (SolarWinds `AlertActive.Acknowledge`\n   / PRTG `acknowledgealarm` / Zabbix `event.acknowledge`) — the prior ack state\n   is captured into priorState, and the ack is double-confirmed\n6. **Failure branch**: if `doctor` fails, do **not** acknowledge anything — you\n   would be silencing alerts you cannot currently see. Fix credentials with\n   `monitoring-aiops secret set <target>` first. If you acknowledged the wrong\n   alert, `monitoring-aiops undo list` → `undo apply <id>` restores the prior\n   ack state.\n\n### 2. \"Which nodes are down and what's saturated?\" (read-only)\n\n1. `noc_rollup` → down / warning counts plus the worst-CPU nodes in a single\n   call, so you do not page through a dashboard\n2. `topn cpu` (also `memory`, `latency`, `packetloss`) → the worst offenders\n   with the measured number\n3. `node_status <node>` → drill into one node; `interface_status` for a\n   suspected link problem, `volume_status` for a filling disk,\n   `application_status` for an app-layer fault\n4. `list_events` → what changed around the time things went bad\n5. `list_unmanaged` → check whether a \"missing\" node is simply unmanaged from a\n   previous maintenance window that was never reverted\n6. **Failure branch**: if a node shows down but is reachable from your shell,\n   the fault is in polling, not the node — check `list_muted` and\n   `list_unmanaged` before escalating to the network team.\n\n### 3. Planned maintenance: suppress noise time-boxed, then restore\n\n1. `node_status <node>` / `swql_canned nodes_down` → confirm you have the right\n   node and that it is currently healthy (so you can tell the difference\n   afterwards)\n2. Prefer the **time-boxed** path — it expires on its own:\n   `schedule_maintenance <node> --end ...` (SolarWinds),\n   `schedule_maintenance_prtg` (PRTG), or `zabbix_create_maintenance` (Zabbix,\n   undo → delete that maintenance id)\n3. If you genuinely need to unmanage instead:\n   `unmanage_node <node> --dry-run`, then re-run without `--dry-run` →\n   **high** risk, double confirmation; it records an inverse `remanage_node`\n   undo descriptor\n4. For a single noisy sensor rather than a whole node: `pause_sensor` (PRTG,\n   undo → `resume_sensor`) or `mute_alerts` (undo → `unmute_alerts`)\n5. When maintenance ends: `remanage_node <node>` / `resume_sensor` /\n   `unmute_alerts`, or simply `monitoring-aiops undo apply <id>` to replay the\n   recorded inverse\n6. **Failure branch**: the classic failure here is *forgetting to restore* —\n   run `list_unmanaged` and `list_muted` at the end of every maintenance window;\n   anything still listed is silently unmonitored. Time-boxed maintenance windows\n   are preferred precisely because they fail safe.\n\n### 4. Answer a bespoke NOC question with SWQL\n\n1. `monitoring-aiops swql library` (MCP: `swql_library`) → the canned queries,\n   so you do not hand-write what already exists\n2. `monitoring-aiops swql canned nodes_down` → run a canned one directly\n   (also `high_cpu_nodes` and the rest of the library)\n3. Not canned? `monitoring-aiops swql query \"SELECT ...\"` → the passthrough\n   **validates the statement is a read-only SELECT** before it runs; anything\n   else is refused\n4. Feed the result into an action — e.g. a node the query surfaced goes into\n   workflow 3 for a maintenance window\n5. **Failure branch**: a rejected query is almost always a non-SELECT statement\n   or a SWQL/SQL dialect slip (SWQL has no `*` expansion on some entities).\n   Start from the nearest canned query in `swql library` and modify it rather\n   than writing from scratch. The passthrough will not be talked into a write —\n   writes go through the governed tools, where they are audited.\n\n## Governance & Safety\n\n- Every tool is audited to `~/.monitoring-aiops/audit.db` (relocatable via\n  `MONITORING_AIOPS_HOME`).\n- Each tool's `risk_level` is carried into the audit row as a descriptive tier\n  (a label, not a gate). `MONITORING_AUDIT_APPROVED_BY` /\n  `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n- Destructive writes support `--dry-run` and double confirmation at the CLI.\n- Suppression / maintenance writes are **time-boxed** (require an end time /\n  duration). Reversible writes record an inverse descriptor (mute→unmute,\n  unmanage→remanage, pause→resume, zabbix_create_maintenance→delete that\n  maintenance id). `zabbix_delete_maintenance` captures the window's full\n  definition into priorState before deleting.\n\n## References\n\n- `references/capabilities.md` — full tool + platform + SWQL/API-path reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, credentials, and connectivity\n\nFile v0.10.4:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"monitoring-aiops\",\n  \"version\": \"0.10.4\",\n  \"publishedAt\": 1789601158731\n}\n\nFile v0.10.4:references/agent-guardrails.md\n\n# Agent guardrails — running monitoring-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The account you connect with.** Give it a SolarWinds/PRTG/Zabbix login with\n  read-only monitoring scope. A write then fails at the server, which is the\n  only place the permission actually lives — no skill-side flag can be argued\n  around by a model, but a revoked permission cannot be.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Never write to the monitoring database\" | `swql_query` accepts a single read-only `SELECT` and nothing else — no verb invoke, no multi-statement, no DELETE. Orion state changes only happen through the named, governed write tools. |\n| \"Don't invent a value when a field is missing\" | A column the platform did not return comes back as `null`, never as `\"\"`. An absent Orion `StatusDescription`, a PRTG sensor `message`, a Zabbix host `dns` — all distinguishable from a genuinely empty one. |\n| \"Tell me if the output was cut off\" | Every row-capped read returns `returned` / `limit` / `truncated`: `swql_query`, `swql_canned`, `list_events`, `zabbix_events`, `zabbix_item_history`, and `interface_status` with a `top`. Truncation is measured — one row past the cap is fetched, or the full set is counted before the cut — never guessed from a length coincidence. |\n| \"Deduplicate the alert storm before showing me\" | `active_alerts` already rolls repeats of the same message into one row with a `count` and up to three `examples`, worst-first. Report the rollup; do not re-count the raw list. |\n| \"Normalise severity across platforms\" | Zabbix's 0–5 scale is already mapped to canonical `level` values (`info`/`warning`/`high`/`critical`) alongside the platform's own `severity` name. Use `level` for cross-platform statements and `severity` when quoting the platform. |\n| \"Confirm before anything disruptive\" | `remove_node`, `unmanage_node`, `mute_alerts` and the maintenance-window writes all take `dry_run=True` for a preview, and `remove_node`, `unmanage_node` and `zabbix_delete_maintenance` are `risk=high`. ⚠️ **The double confirmation is a CLI feature, and of the writes only `alert_acknowledge` and `undo apply` have CLI commands** — every tool named here is reachable only over MCP, where nothing prompts. Keep your own confirmation for them. |\n| \"Log what you did\" | Every call is audited to `~/.monitoring-aiops/audit.db` regardless of what the model says it did. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate an enterprise monitoring system through the monitoring-aiops MCP\ntools. A target is SolarWinds Orion, Paessler PRTG, or Zabbix — check which\nbefore reasoning about what a field means.\n\nTOOL USE\n- Before answering any question about the current monitored estate, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"everything is fine\".\n- Prefer the named canned SWQL (swql_library / swql_canned) over writing your\n  own query. Hand-written SWQL is where small models most often produce\n  syntactically valid nonsense.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete. An alert count from a truncated read is not\n  the number of alerts.\n- A null field means the platform did not return that value. Report it as \"not\n  available\" — never infer it.\n- Report values exactly as returned. Do not translate status codes, severity\n  names, node captions, or sensor names into your own vocabulary.\n- Acknowledged is not the same as resolved. An acknowledged alert is still\n  active; say which you mean.\n\n- None of the node or maintenance-window writes has a CLI command, so nothing will ask you\n  to confirm them — `remove_node` and `unmanage_node` least of all. Call with\n  `dry_run=True` first and wait for an explicit go-ahead.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, root cause, or business impact unless a tool result\n  supports it. A down sensor is one sensor, not necessarily a down service.\n- Do not confuse the identifier kinds: an Orion node Caption, an AlertActiveID,\n  a PRTG objid, and a Zabbix eventid/triggerid/itemid are different namespaces\n  and are not interchangeable across targets.\n- Muting, unmanaging, and maintenance windows suppress alerting; they do not fix\n  anything. Never describe them as a resolution.\n```\n\n## Recommended setup for a local model\n\nStart with a connection that *cannot* write — a SolarWinds/PRTG/Zabbix account\nwith read-only monitoring scope — verify, and widen the account's permission\nonly when you trust the setup:\n\n```bash\nmonitoring-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport MONITORING_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport MONITORING_AUDIT_RATIONALE=\"change window CHG0041231, muting core-sw1\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Lead with `noc_rollup` (SolarWinds)\n  or `active_alerts` — they do the correlation and dedup inside one call, so the\n  model does not have to chain reads and keep ids straight.\n- **The model writes broken SWQL.** Use `swql_library` and `swql_canned` — the\n  canned queries answer the most-asked questions and are already parameterised.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions and use `top` / `limit` deliberately rather than pulling every\n  sensor in the estate.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.4:references/capabilities.md\n\n# monitoring-aiops capabilities\n\n> **42 MCP tools** (30 read, 10 write, 2 undo) across SolarWinds\n> Orion (SWIS REST + SWQL, port 17774 with a legacy-17778 fallback, HTTP\n> Basic auth), Paessler PRTG (web\n> API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at\n> `/api_jsonrpc.php`, API token — Bearer header on 6.4+/7.x, legacy `auth`\n> field fallback for 6.0). Each config target names its own `platform`.\n> SWIS/PRTG/Zabbix responses are mocked and need live verification.\n\n## SWQL — SolarWinds (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `swql_library` | (local) | the catalogue of canned queries: `nodes_down`, `flapping_interfaces`, `muted_report`, `high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled` |\n| `swql_canned` | named SWQL → SWIS `/Query` | rows for the named canned query |\n| `swql_query` | validated read-only SWQL → SWIS `/Query` | rows for a caller SELECT (SELECT-only; rejected otherwise) |\n\n## Alerts — all platforms\n\n| Tool | Risk | Path | Returns / effect |\n|------|------|------|------------------|\n| `active_alerts` | read | SWIS `AlertActive`/`AlertObjects`, PRTG `/api/table.json?content=messages`, or Zabbix `problem.get` | active alerts **deduped/rolled up by message** — flap/down storms collapse into one counted entry |\n| `alert_acknowledge` | write **medium** | SW `AlertActive.Acknowledge` verb / PRTG `acknowledgealarm.htm` / Zabbix `event.acknowledge` (action 6; prior ack state → priorState) | acknowledges an alert / alarm / problem event |\n\n## SolarWinds health (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `node_status` | `Orion.Nodes` | one node's status, CPU/mem, response time |\n| `nodes_list` | `Orion.Nodes` | node inventory (status, vendor, IP, last boot) |\n| `interface_status` | `Orion.NPM.Interfaces` | top-N interfaces by utilisation (in/out, errors, oper status) |\n| `volume_status` | `Orion.Volumes` | volumes by % used (size, used, type) |\n| `application_status` | `Orion.APM.Application` (SAM) | SAM application/component status |\n| `topn` | `Orion.Nodes` metrics | top-N nodes by `cpu` / `memory` / `latency` / `packetloss` |\n| `noc_rollup` | folds `Orion.Nodes` | down/warning counts + worst-CPU nodes in one call |\n\n## SolarWinds writes\n\n| Tool | Risk | Path / verb | Undo / safety |\n|------|------|-------------|---------------|\n| `list_events` | read | `Orion.Events` | recent events (read) |\n| `list_unmanaged` | read | `Orion.Nodes` where Unmanaged | currently-unmanaged nodes (read) |\n| `list_muted` | read | `Orion.AlertSuppression` | currently-muted objects (read) |\n| `mute_alerts` | write **med** | `AlertSuppression` (SuppressAlerts) | **time-boxed** (requires end time); records inverse **unmute** undo |\n| `unmute_alerts` | write **med** | `AlertSuppression` (ResumeAlerts) | un-suppresses alerting |\n| `schedule_maintenance` | write **med** | `AlertSuppression` window | **requires an end time** (time-boxed maintenance window) |\n| `unmanage_node` | write **HIGH** | `Orion.Nodes.Unmanage` verb | `dry_run` + double-confirm; captures prior managed state; records inverse **remanage** undo |\n| `remanage_node` | write **med** | `Orion.Nodes.Remanage` verb | brings a node back under management |\n| `remove_node` | write **HIGH** | SWIS `DELETE` on the node URI | `dry_run` + double-confirm; no undo (deletion is not reversible) |\n\n## PRTG (read)\n\n| Tool | Path | Returns |\n|------|------|---------|\n| `prtg_sensors` | `/api/table.json?content=sensors` | sensors (status, last value, message) |\n| `prtg_sensor_details` | `/api/getsensordetails.json` | one sensor's detail (channels, uptime, last check) |\n| `prtg_devices` | `/api/table.json?content=devices` | devices (host, group, status) |\n| `prtg_groups` | `/api/table.json?content=groups` | probe/group tree with status rollups |\n| `prtg_history` | `/api/historicdata.json` | historic values for a sensor over a window |\n| `prtg_system_status` | `/api/status.json` | server/system status (also the PRTG `doctor` check) |\n| `prtg_alarms` | `/api/table.json?content=messages` (alarms) | active PRTG alarms |\n\n## PRTG writes\n\n| Tool | Risk | Path | Undo / safety |\n|------|------|------|---------------|\n| `pause_sensor` | write **med** | `/api/pause.htm?action=0` | records inverse **resume** undo |\n| `resume_sensor` | write **med** | `/api/pause.htm?action=1` | resumes a paused sensor |\n| `schedule_maintenance_prtg` | write **med** | `/api/pauseobjectfor.htm?duration=` | **time-boxed** (requires minutes) |\n\n## Zabbix (read)\n\n| Tool | JSON-RPC method | Returns |\n|------|-----------------|---------|\n| `zabbix_problems` | `problem.get` | current problems; severity 0-5 mapped to names + canonical levels (`info`/`warning`/`high`/`critical`) |\n| `zabbix_hosts` | `host.get` (+interfaces, +host groups) | host inventory: monitored flag, interfaces (ip/dns/availability), groups |\n| `zabbix_hostgroups` | `hostgroup.get` | host groups (ids + names) |\n| `zabbix_triggers` | `trigger.get` | triggers, by default only those currently firing (PROBLEM) |\n| `zabbix_events` | `event.get` | recent trigger events (newest first, capped) |\n| `zabbix_item_history` | `item.get` + `history.get` | **bounded** metric detail: item meta + history points (window ≤ 168 h, ≤ 500 points) |\n| `zabbix_maintenances` | `maintenance.get` | maintenance windows with hosts/groups + periods |\n\n## Zabbix writes\n\n| Tool | Risk | JSON-RPC method | Undo / safety |\n|------|------|-----------------|---------------|\n| `zabbix_create_maintenance` | write **med** | `maintenance.create` | **time-boxed** (minutes > 0) + must name hosts/groups; records a **replayable undo** = delete exactly the created maintenance id |\n| `zabbix_delete_maintenance` | write **HIGH** | `maintenance.delete` | `dry_run` + double-confirm; the window's **full definition** is captured into priorState first (no undo — re-create manually from it) |\n\n## Out of scope (by design)\n\n- Monitoring stacks other than SolarWinds Orion, PRTG, and Zabbix\n- Zabbix template / discovery / user CRUD, `trend.get`, and host **onboarding**\n- Creating alerts / thresholds / SWQL-view CRUD, and node/interface **onboarding**\n- Anything outside monitoring (hypervisor, storage, backup, cluster, network\n  device config, OT/industrial) — route to the appropriate other AIops-tools skill\n\nWant one of these? Open an issue or PR — feedback and contributions welcome.\n\nFile v0.10.4:references/cli-reference.md\n\n# monitoring-aiops CLI reference\n\n> Covers SolarWinds Orion (SWIS REST + SWQL), Paessler\n> PRTG (web API), and Zabbix 6.x/7.x (JSON-RPC); SWIS/PRTG/Zabbix responses are\n> mocked and need live verification.\n> The CLI is a convenience subset — the full 42-tool surface is via the MCP\n> server (`monitoring-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nmonitoring-aiops init                      # interactive wizard (asks for the platform: solarwinds/prtg/zabbix)\nmonitoring-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   SolarWinds: a SWQL query · PRTG: /api/status.json\n                                           #   Zabbix: apiinfo.version (no auth) + authed host count\nmonitoring-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.monitoring-aiops/secrets.enc)\n\n```bash\nmonitoring-aiops secret set <target> [--value <secret>]  # store Orion password / PRTG or Zabbix token (hidden prompt if no --value)\nmonitoring-aiops secret list                             # names only — secrets never shown\nmonitoring-aiops secret rm <target>\nmonitoring-aiops secret migrate                          # import legacy plaintext env (MONITORING_<TARGET>_SECRET)\nmonitoring-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nmonitoring-aiops overview [--target <t>]   # NOC summary: platform + active/unacked alert counts + top rollup\n```\n\n## SWQL (SolarWinds)\n\n```bash\nmonitoring-aiops swql library                    # list the canned queries\nmonitoring-aiops swql canned <name>              # run a canned query: nodes_down, flapping_interfaces,\n                                                 #   muted_report, high_cpu_nodes, volumes_full, unmanaged_scheduled\nmonitoring-aiops swql query \"SELECT ...\"         # validated read-only SWQL passthrough (SELECT only)\n```\n\n## Alerts (all platforms)\n\n```bash\nmonitoring-aiops alert list [--target <t>]       # active alerts, deduped/rolled up by message\nmonitoring-aiops alert ack <alert_id>            # acknowledge an alert / PRTG alarm / Zabbix problem event\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `swql`, and `alert` are the CLI subset; the remaining SolarWinds\n  health, PRTG, Zabbix, and governed-write tools (mute/unmute,\n  schedule_maintenance, unmanage/remanage/remove node, PRTG pause/resume,\n  Zabbix maintenance create/delete) are exposed through the MCP server.\n  High-risk MCP writes use dry-run + double-confirm; `MONITORING_AUDIT_APPROVED_BY`\n  / `MONITORING_AUDIT_RATIONALE` are recorded on the audit row when set.\n\nFile v0.10.4:references/setup-guide.md\n\n# monitoring-aiops setup & security guide\n\n> Not yet validated against a live NOC (see `docs/VERIFICATION.md`). **PRTG's free\n> perpetual 100-sensor Freeware edition (with the API) and an open-source\n> Zabbix appliance (Docker compose) are the easiest live checks; SolarWinds is\n> a 30-day trial only — mock-only past that, the largest verification debt.**\n\n## 1. Install\n\n```bash\nuv tool install monitoring-aiops\n```\n\n## 2. Get a credential\n\n- **SolarWinds Orion** — an Orion account (username + password) with the Orion\n  API enabled. monitoring-aiops talks to SWIS REST + SWQL with **HTTP Basic\n  auth** on port **17774** — the SWIS port since Orion 2023.1, which deprecated\n  the old **17778** and slated it for removal. Leave `port` unset and a\n  pre-2023.1 server is still reached: the first connection failure triggers one\n  retry on 17778, and the port that answered is reused for that session. Set\n  `port:` yourself and it is used verbatim, with no fallback probing.\n- **PRTG** — an **API token** (Setup → Account Settings → API Keys, or a\n  passhash). PRTG's web API is on port **443/8080**. A free Freeware edition\n  (100 sensors, perpetual) exposes the same API — the easiest way to self-test.\n- **Zabbix (6.x/7.x)** — an **API token** (Administration → API tokens; user\n  tokens under User settings → API tokens). monitoring-aiops talks JSON-RPC 2.0\n  to `/api_jsonrpc.php` on the frontend port (default **443**), sending the\n  token as a `Bearer` header (6.4+/7.x) with an automatic legacy `auth`-field\n  fallback for 6.0. Zabbix is fully open source — a Docker-compose appliance\n  is a 10-minute self-test.\n\n## 3. Onboard\n\n```bash\nmonitoring-aiops init\n```\n\nThe wizard asks, per target, for the **platform** (`solarwinds` / `prtg` /\n`zabbix`), the **host**, the **port** (defaults 17774 for SolarWinds, 443 for\nPRTG and Zabbix — accept the default and no `port:` key is written, so the\ndefault keeps tracking the platform), the **Orion username** (SolarWinds only),\nand the **secret**\n— the Orion account password, the PRTG API token, or the Zabbix API token.\nNon-secret connection details go to `~/.monitoring-aiops/config.yaml`; the\nsecret is stored **encrypted** into `~/.monitoring-aiops/secrets.enc`. Example\nconfig (one config can span all NOCs):\n\n```yaml\ntargets:\n  - name: orion1\n    platform: solarwinds\n    host: 10.0.0.20\n    # port omitted -> 17774 (Orion 2023.1+), falling back to 17778 once if\n    # nothing answers. Set it here only to pin a non-standard SWIS port.\n    username: admin\n    verify_ssl: false          # self-signed lab certs only\n  - name: prtg1\n    platform: prtg\n    host: 10.0.0.40\n    port: 443\n    verify_ssl: true\n  - name: zbx1\n    platform: zabbix\n    host: 10.0.0.60\n    port: 443\n    verify_ssl: true\n```\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nExport the master password so the encrypted store can be unlocked without a\nprompt:\n\n```bash\nexport MONITORING_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## Credential security\n\n- The secret (Orion password / PRTG API token / Zabbix API token) is **never**\n  written to disk in\n  plaintext. It lives only in `~/.monitoring-aiops/secrets.enc`, encrypted with\n  Fernet (AES-128-CBC + HMAC), the key derived from your master password via\n  scrypt. Only a per-store random salt and the ciphertext are on disk (chmod\n  600); the master password itself is never stored.\n- A legacy plaintext env var `MONITORING_<TARGET_NAME_UPPER>_SECRET` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `monitoring-aiops secret migrate` (it imports then renames the old `.env`).\n- The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG / Zabbix\n  API token at request time and held only in memory; it is never logged or\n  echoed. Exception\n  text and tracebacks are scrubbed of secret-shaped strings before being written\n  to the audit log.\n\n## Governance harness state\n\nState lives under `~/.monitoring-aiops/` (relocate with `MONITORING_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with risk tier, approver, rationale\n- `undo.db` — inverse descriptors for reversible writes (mute→unmute,\n  unmanage→remanage, pause→resume, zabbix create-maintenance→delete)\n- budget / runaway guard — caps cumulative tool calls and wall-time; trips on\n  tight poll/retry loops\n\n## Governed writes\n\n- **High-risk** ops (`unmanage_node`, `remove_node`,\n  `zabbix_delete_maintenance`) use `dry_run` + double confirmation;\n  `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional\n  audit annotations, recorded when set but never required.\n- **Time-boxed** ops require an end time / duration: `mute_alerts`,\n  `schedule_maintenance` (SolarWinds), `schedule_maintenance_prtg` (PRTG, in\n  minutes), and `zabbix_create_maintenance` (Zabbix, in minutes). This prevents\n  forgotten, indefinite suppression windows.\n\n## Verify\n\n```bash\nmonitoring-aiops doctor\n```\n\n`doctor` is platform-aware: it checks the config file, the encrypted store and\nits permissions, that a secret is present per target, and (unless `--skip-auth`)\nconnectivity — a SWQL query for SolarWinds targets, `/api/status.json` for PRTG\ntargets, and for Zabbix targets the unauthenticated `apiinfo.version`\n(reachability) followed by a cheap authed host count (token validity).\n\nFile v0.10.4:skill-card.md\n\n## Description:\n\nMonitoring AIops lets agents operate SolarWinds Orion, Paessler PRTG, and Zabbix monitoring environments for NOC overviews, alert rollups, read-only SWQL queries, health checks, and governed write actions.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, SREs, and NOC operators use this skill to inspect and triage infrastructure monitoring data across SolarWinds, PRTG, and Zabbix, then perform audited maintenance or alert actions when explicitly authorized.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill exposes high-impact monitoring write actions without an enforced MCP approval or read-only gate.\n\nMitigation: Start with read-only SolarWinds, PRTG, or Zabbix credentials and require explicit human approval before MCP writes such as mute, unmanage, remove node, pause sensor, or delete maintenance.\n\nRisk: Production monitoring access can expose sensitive credentials and operational state.\n\nMitigation: Use the encrypted credential setup, set the master password only through controlled environment configuration, and enable SSL verification for production targets.\n\nRisk: Some behavior is documented as exercised against mocked monitoring responses rather than live NOC systems.\n\nMitigation: Run connectivity and behavior checks with `monitoring-aiops doctor` and validate workflows against low-impact PRTG or Zabbix environments before using production SolarWinds workflows.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/monitoring-aiops)\n- [Monitoring AIops homepage](https://github.com/AIops-tools/Monitoring-AIops)\n- [Capabilities reference](references/capabilities.md)\n- [CLI reference](references/cli-reference.md)\n- [Setup and security guide](references/setup-guide.md)\n- [Agent guardrails](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and operational guidance]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May direct an agent to call MCP tools or CLI commands against configured SolarWinds, PRTG, or Zabbix targets; read results may be truncated and should be reported as such.]\n\n## Skill Version(s):\n\n0.10.4 (source: ClawHub release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.3: 7 files, 19423 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2968b), SKILL.md (16987b), _meta.json (136b)\n\nFile v0.10.3:SKILL.md\n\n---\nname: monitoring-aiops\nslug: monitoring-aiops\ndisplayName: \"Monitoring AIops\"\nsummary: \"Governed SolarWinds Orion + PRTG + Zabbix ops: SWQL, alert rollup, health, 42 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Monitoring-AIops\ntags: [aiops, mcp, governance, monitoring]\ndescription: >\n  Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window).\n  Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring.\n  Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill.\n  Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.\ninstaller:\n  kind: uv\n  package: monitoring-aiops\nargument-hint: \"[node/sensor id, a SWQL question, or describe your NOC task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"monitoring-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"MONITORING_AIOPS_CONFIG\",\"MONITORING_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Monitoring-AIops\",\"emoji\":\"📡\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed monitoring operations across SolarWinds Orion (SWIS REST + SWQL, port 17774 on Orion 2023.1+ with an automatic one-shot fallback to the legacy 17778, HTTP Basic auth), Paessler PRTG (web API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at /api_jsonrpc.php, API token as Bearer header on 6.4+/7.x with a legacy auth-field fallback for 6.0). Each target in the config names its own platform, so one config can span all NOCs. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.monitoring-aiops/ (relocatable via MONITORING_AIOPS_HOME).\n  Credentials: the Orion account password (SolarWinds), the PRTG API token, or the Zabbix API token is stored ENCRYPTED in ~/.monitoring-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'monitoring-aiops init' to onboard (it asks for the platform), or 'monitoring-aiops secret set <target>' to add one. The store is unlocked by a master password from MONITORING_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var MONITORING_<TARGET_NAME_UPPER>_SECRET is still honoured as a fallback with a deprecation warning (migrate with 'monitoring-aiops secret migrate'). The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG/Zabbix API token at request time and held only in memory; secrets are never logged or echoed.\n  Read-only SWQL passthrough (swql_query) is validated to accept SELECT statements only. State-changing operations pass through the @governed_tool decorator (budget guard + audit + undo recording; each tool's risk_level is recorded as a descriptive tier, not a gate). Destructive writes (unmanage_node, remove_node, zabbix_delete_maintenance) are high-risk with dry_run + double confirmation; unmanage_node records an inverse remanage undo descriptor, zabbix_delete_maintenance captures the window's FULL definition into priorState first. Suppression/maintenance writes are TIME-BOXED (mute_alerts, schedule_maintenance, schedule_maintenance_prtg, zabbix_create_maintenance require an end time / duration). mute_alerts→unmute, pause_sensor→resume, zabbix_create_maintenance→delete-that-maintenance-id record inverse undo descriptors. Zabbix item history is BOUNDED (capped window + point count).\n  Webhooks: none — no outbound network calls beyond the configured SolarWinds SWIS / PRTG web API / Zabbix JSON-RPC endpoint.\n  SSL: verify_ssl defaults to false-friendly for self-signed lab certs; enable for production.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked SWIS/PRTG/Zabbix responses; not yet run against a live NOC (see docs/VERIFICATION.md). PRTG has a free perpetual 100-sensor Freeware edition with the API, and Zabbix is fully open source (a Docker-compose appliance is a 10-minute live check) — the easiest live checks; SolarWinds is a 30-day trial (mock-only past that — largest verification debt).\n---\n\n# Monitoring AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by SolarWinds, Paessler, Zabbix, or any monitoring vendor.** SolarWinds, Orion, SWQL, THWACK, PRTG, Paessler and Zabbix are trademarks of their respective owners. Source at [github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops) under the MIT license.\n\nGoverned network / infrastructure monitoring operations — **42 MCP tools**\nacross **SolarWinds Orion** (SWIS REST + SWQL), **Paessler PRTG** (web API),\nand **Zabbix 6.x/7.x** (JSON-RPC 2.0),\nevery one wrapped with the bundled `@governed_tool` harness: a local unified\naudit log under `~/.monitoring-aiops/`, policy engine, token/runaway budget\nguard, undo-token recording, and risk-tier labelling on the audit trail. One\nconfig can span all NOCs. The Orion password / PRTG API token / Zabbix API token is stored\n**encrypted** (`~/.monitoring-aiops/secrets.enc`, Fernet + scrypt) — never\nplaintext on disk.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`monitoring_aiops.governance`) — no external skill-family dependency.\n> PRTG's free Freeware edition and an open-source Zabbix appliance are the\n> easiest live checks; SolarWinds is trial-only past 30 days (largest\n> verification debt — see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **SWQL** | SolarWinds | library, canned, query (SELECT-only passthrough) | 3 | read |\n| **Alerts** | all | active_alerts (dedup/rollup), alert_acknowledge | 2 | 1 read, 1 write |\n| **SolarWinds health** | SolarWinds | node/nodes/interface/volume/application status, topn, noc_rollup | 7 | read |\n| **SolarWinds writes** | SolarWinds | list_events/unmanaged/muted | 3 | read |\n| | SolarWinds | mute/unmute, schedule_maintenance, remanage_node | 4 | write (med) |\n| | SolarWinds | unmanage_node, remove_node | 2 | write (**high**) |\n| **PRTG** | PRTG | sensors/sensor_details/devices/groups/history/system_status/alarms | 7 | read |\n| **PRTG writes** | PRTG | pause_sensor, resume_sensor, schedule_maintenance_prtg | 3 | write (med) |\n| **Zabbix** | Zabbix | zabbix_problems/hosts/hostgroups/triggers/events/item_history/maintenances | 7 | read |\n| **Zabbix writes** | Zabbix | zabbix_create_maintenance (time-boxed; undo = delete that id) | 1 | write (med) |\n| | Zabbix | zabbix_delete_maintenance (priorState = full definition) | 1 | write (**high**) |\n| **Undo** | all | undo_list, undo_apply | 2 | undo |\n\nThe canned SWQL library (`swql_library` lists them) answers the most-repeated\nTHWACK questions directly: `nodes_down`, `flapping_interfaces`, `muted_report`,\n`high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled`. For anything else,\n`swql_query` is a validated read-only (SELECT-only) SWQL passthrough.\n\n## Quick Install\n\n```bash\nuv tool install monitoring-aiops\nmonitoring-aiops init       # wizard: pick platform (solarwinds/prtg/zabbix) + encrypted secret\nmonitoring-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/monitoring-aiops\nopenclaw skills info monitoring-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a NOC snapshot (`overview` / `noc_rollup`): active/unacked alert counts,\n  down/warning nodes, worst CPU\n- Answer a repeated SWQL question (`swql_library` → `swql_canned nodes_down`), or\n  run an ad-hoc read-only SWQL SELECT (`swql_query`)\n- Triage an alert storm (`active_alerts` dedup/rollup collapses flap/down\n  storms), then `alert_acknowledge`\n- SolarWinds health: `node_status`, `interface_status` (top-N by util),\n  `volume_status`, `application_status` (SAM), `topn` (cpu/mem/latency/loss)\n- PRTG: list `prtg_sensors` / `prtg_devices` / `prtg_groups`, drill with\n  `prtg_sensor_details` / `prtg_history`, check `prtg_alarms` / `prtg_system_status`\n- Zabbix: triage `zabbix_problems` (0-5 severity mapped to levels) /\n  `zabbix_triggers`, inventory `zabbix_hosts` / `zabbix_hostgroups`, drill with\n  `zabbix_item_history` (bounded), review `zabbix_events` / `zabbix_maintenances`\n- Safely take a node out for maintenance (`schedule_maintenance` /\n  `unmanage_node` with dry_run + double-confirm), pause a PRTG sensor\n  (`pause_sensor`), or create a time-boxed Zabbix maintenance window\n  (`zabbix_create_maintenance` — undo deletes exactly that window)\n\n**Do NOT use when** the target is not a SolarWinds/PRTG/Zabbix monitoring\nplatform — route hypervisor, storage, backup, cluster, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| SolarWinds Orion / SWQL, PRTG, or Zabbix monitoring ops | **monitoring-aiops** (this skill) |\n| A non-monitoring platform (hypervisor, storage, backup, cluster, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Other monitoring stacks (not SolarWinds/PRTG/Zabbix) | out of scope for this tool |\n\n## Common Workflows\n\n> **No authorization gate**: the skill runs the operations you ask for and audits every one; it does not decide whether a write is permitted — that is the agent's judgement or the permissions of the SolarWinds/PRTG/Zabbix account it connects with (a read-only monitoring account makes writes fail at the server). There is no read-only switch, policy file, or approval gate. `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n\n### 1. The 3 a.m. alert storm — collapse it, then acknowledge what matters\n\n1. `monitoring-aiops doctor` → confirm the NOC platform is actually reachable\n   (a \"storm\" is sometimes just a poller that lost the target)\n2. `monitoring-aiops overview` → the one-screen picture: down/warning counts\n   across the configured targets\n3. `monitoring-aiops alert list` (MCP: `active_alerts`) → deduped / rolled-up\n   entries; an interface-flap or node-down storm collapses into **one** entry\n   with a count instead of a wall of alerts\n4. `noc_rollup` → confirm whether the storm has a single upstream cause (one\n   node down taking its children with it) rather than N independent faults\n5. Acknowledge only the rolled-up entry that matters:\n   `monitoring-aiops alert ack <alert-id>` (SolarWinds `AlertActive.Acknowledge`\n   / PRTG `acknowledgealarm` / Zabbix `event.acknowledge`) — the prior ack state\n   is captured into priorState, and the ack is double-confirmed\n6. **Failure branch**: if `doctor` fails, do **not** acknowledge anything — you\n   would be silencing alerts you cannot currently see. Fix credentials with\n   `monitoring-aiops secret set <target>` first. If you acknowledged the wrong\n   alert, `monitoring-aiops undo list` → `undo apply <id>` restores the prior\n   ack state.\n\n### 2. \"Which nodes are down and what's saturated?\" (read-only)\n\n1. `noc_rollup` → down / warning counts plus the worst-CPU nodes in a single\n   call, so you do not page through a dashboard\n2. `topn cpu` (also `memory`, `latency`, `packetloss`) → the worst offenders\n   with the measured number\n3. `node_status <node>` → drill into one node; `interface_status` for a\n   suspected link problem, `volume_status` for a filling disk,\n   `application_status` for an app-layer fault\n4. `list_events` → what changed around the time things went bad\n5. `list_unmanaged` → check whether a \"missing\" node is simply unmanaged from a\n   previous maintenance window that was never reverted\n6. **Failure branch**: if a node shows down but is reachable from your shell,\n   the fault is in polling, not the node — check `list_muted` and\n   `list_unmanaged` before escalating to the network team.\n\n### 3. Planned maintenance: suppress noise time-boxed, then restore\n\n1. `node_status <node>` / `swql_canned nodes_down` → confirm you have the right\n   node and that it is currently healthy (so you can tell the difference\n   afterwards)\n2. Prefer the **time-boxed** path — it expires on its own:\n   `schedule_maintenance <node> --end ...` (SolarWinds),\n   `schedule_maintenance_prtg` (PRTG), or `zabbix_create_maintenance` (Zabbix,\n   undo → delete that maintenance id)\n3. If you genuinely need to unmanage instead:\n   `unmanage_node <node> --dry-run`, then re-run without `--dry-run` →\n   **high** risk, double confirmation; it records an inverse `remanage_node`\n   undo descriptor\n4. For a single noisy sensor rather than a whole node: `pause_sensor` (PRTG,\n   undo → `resume_sensor`) or `mute_alerts` (undo → `unmute_alerts`)\n5. When maintenance ends: `remanage_node <node>` / `resume_sensor` /\n   `unmute_alerts`, or simply `monitoring-aiops undo apply <id>` to replay the\n   recorded inverse\n6. **Failure branch**: the classic failure here is *forgetting to restore* —\n   run `list_unmanaged` and `list_muted` at the end of every maintenance window;\n   anything still listed is silently unmonitored. Time-boxed maintenance windows\n   are preferred precisely because they fail safe.\n\n### 4. Answer a bespoke NOC question with SWQL\n\n1. `monitoring-aiops swql library` (MCP: `swql_library`) → the canned queries,\n   so you do not hand-write what already exists\n2. `monitoring-aiops swql canned nodes_down` → run a canned one directly\n   (also `high_cpu_nodes` and the rest of the library)\n3. Not canned? `monitoring-aiops swql query \"SELECT ...\"` → the passthrough\n   **validates the statement is a read-only SELECT** before it runs; anything\n   else is refused\n4. Feed the result into an action — e.g. a node the query surfaced goes into\n   workflow 3 for a maintenance window\n5. **Failure branch**: a rejected query is almost always a non-SELECT statement\n   or a SWQL/SQL dialect slip (SWQL has no `*` expansion on some entities).\n   Start from the nearest canned query in `swql library` and modify it rather\n   than writing from scratch. The passthrough will not be talked into a write —\n   writes go through the governed tools, where they are audited.\n\n## Governance & Safety\n\n- Every tool is audited to `~/.monitoring-aiops/audit.db` (relocatable via\n  `MONITORING_AIOPS_HOME`).\n- Each tool's `risk_level` is carried into the audit row as a descriptive tier\n  (a label, not a gate). `MONITORING_AUDIT_APPROVED_BY` /\n  `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n- Destructive writes support `--dry-run` and double confirmation at the CLI.\n- Suppression / maintenance writes are **time-boxed** (require an end time /\n  duration). Reversible writes record an inverse descriptor (mute→unmute,\n  unmanage→remanage, pause→resume, zabbix_create_maintenance→delete that\n  maintenance id). `zabbix_delete_maintenance` captures the window's full\n  definition into priorState before deleting.\n\n## References\n\n- `references/capabilities.md` — full tool + platform + SWQL/API-path reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, credentials, and connectivity\n\nFile v0.10.3:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"monitoring-aiops\",\n  \"version\": \"0.10.3\",\n  \"publishedAt\": 1789452361725\n}\n\nFile v0.10.3:references/agent-guardrails.md\n\n# Agent guardrails — running monitoring-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The account you connect with.** Give it a SolarWinds/PRTG/Zabbix login with\n  read-only monitoring scope. A write then fails at the server, which is the\n  only place the permission actually lives — no skill-side flag can be argued\n  around by a model, but a revoked permission cannot be.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Never write to the monitoring database\" | `swql_query` accepts a single read-only `SELECT` and nothing else — no verb invoke, no multi-statement, no DELETE. Orion state changes only happen through the named, governed write tools. |\n| \"Don't invent a value when a field is missing\" | A column the platform did not return comes back as `null`, never as `\"\"`. An absent Orion `StatusDescription`, a PRTG sensor `message`, a Zabbix host `dns` — all distinguishable from a genuinely empty one. |\n| \"Tell me if the output was cut off\" | Every row-capped read returns `returned` / `limit` / `truncated`: `swql_query`, `swql_canned`, `list_events`, `zabbix_events`, `zabbix_item_history`, and `interface_status` with a `top`. Truncation is measured — one row past the cap is fetched, or the full set is counted before the cut — never guessed from a length coincidence. |\n| \"Deduplicate the alert storm before showing me\" | `active_alerts` already rolls repeats of the same message into one row with a `count` and up to three `examples`, worst-first. Report the rollup; do not re-count the raw list. |\n| \"Normalise severity across platforms\" | Zabbix's 0–5 scale is already mapped to canonical `level` values (`info`/`warning`/`high`/`critical`) alongside the platform's own `severity` name. Use `level` for cross-platform statements and `severity` when quoting the platform. |\n| \"Confirm before anything disruptive\" | `remove_node`, `unmanage_node`, `mute_alerts`, and the maintenance-window writes require a `--dry-run`-able preview + double confirmation at the CLI. |\n| \"Log what you did\" | Every call is audited to `~/.monitoring-aiops/audit.db` regardless of what the model says it did. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate an enterprise monitoring system through the monitoring-aiops MCP\ntools. A target is SolarWinds Orion, Paessler PRTG, or Zabbix — check which\nbefore reasoning about what a field means.\n\nTOOL USE\n- Before answering any question about the current monitored estate, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"everything is fine\".\n- Prefer the named canned SWQL (swql_library / swql_canned) over writing your\n  own query. Hand-written SWQL is where small models most often produce\n  syntactically valid nonsense.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete. An alert count from a truncated read is not\n  the number of alerts.\n- A null field means the platform did not return that value. Report it as \"not\n  available\" — never infer it.\n- Report values exactly as returned. Do not translate status codes, severity\n  names, node captions, or sensor names into your own vocabulary.\n- Acknowledged is not the same as resolved. An acknowledged alert is still\n  active; say which you mean.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, root cause, or business impact unless a tool result\n  supports it. A down sensor is one sensor, not necessarily a down service.\n- Do not confuse the identifier kinds: an Orion node Caption, an AlertActiveID,\n  a PRTG objid, and a Zabbix eventid/triggerid/itemid are different namespaces\n  and are not interchangeable across targets.\n- Muting, unmanaging, and maintenance windows suppress alerting; they do not fix\n  anything. Never describe them as a resolution.\n```\n\n## Recommended setup for a local model\n\nStart with a connection that *cannot* write — a SolarWinds/PRTG/Zabbix account\nwith read-only monitoring scope — verify, and widen the account's permission\nonly when you trust the setup:\n\n```bash\nmonitoring-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport MONITORING_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport MONITORING_AUDIT_RATIONALE=\"change window CHG0041231, muting core-sw1\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Lead with `noc_rollup` (SolarWinds)\n  or `active_alerts` — they do the correlation and dedup inside one call, so the\n  model does not have to chain reads and keep ids straight.\n- **The model writes broken SWQL.** Use `swql_library` and `swql_canned` — the\n  canned queries answer the most-asked questions and are already parameterised.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions and use `top` / `limit` deliberately rather than pulling every\n  sensor in the estate.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.3:references/capabilities.md\n\n# monitoring-aiops capabilities\n\n> **42 MCP tools** (30 read, 10 write, 2 undo) across SolarWinds\n> Orion (SWIS REST + SWQL, port 17774 with a legacy-17778 fallback, HTTP\n> Basic auth), Paessler PRTG (web\n> API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at\n> `/api_jsonrpc.php`, API token — Bearer header on 6.4+/7.x, legacy `auth`\n> field fallback for 6.0). Each config target names its own `platform`.\n> SWIS/PRTG/Zabbix responses are mocked and need live verification.\n\n## SWQL — SolarWinds (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `swql_library` | (local) | the catalogue of canned queries: `nodes_down`, `flapping_interfaces`, `muted_report`, `high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled` |\n| `swql_canned` | named SWQL → SWIS `/Query` | rows for the named canned query |\n| `swql_query` | validated read-only SWQL → SWIS `/Query` | rows for a caller SELECT (SELECT-only; rejected otherwise) |\n\n## Alerts — all platforms\n\n| Tool | Risk | Path | Returns / effect |\n|------|------|------|------------------|\n| `active_alerts` | read | SWIS `AlertActive`/`AlertObjects`, PRTG `/api/table.json?content=messages`, or Zabbix `problem.get` | active alerts **deduped/rolled up by message** — flap/down storms collapse into one counted entry |\n| `alert_acknowledge` | write **medium** | SW `AlertActive.Acknowledge` verb / PRTG `acknowledgealarm.htm` / Zabbix `event.acknowledge` (action 6; prior ack state → priorState) | acknowledges an alert / alarm / problem event |\n\n## SolarWinds health (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `node_status` | `Orion.Nodes` | one node's status, CPU/mem, response time |\n| `nodes_list` | `Orion.Nodes` | node inventory (status, vendor, IP, last boot) |\n| `interface_status` | `Orion.NPM.Interfaces` | top-N interfaces by utilisation (in/out, errors, oper status) |\n| `volume_status` | `Orion.Volumes` | volumes by % used (size, used, type) |\n| `application_status` | `Orion.APM.Application` (SAM) | SAM application/component status |\n| `topn` | `Orion.Nodes` metrics | top-N nodes by `cpu` / `memory` / `latency` / `packetloss` |\n| `noc_rollup` | folds `Orion.Nodes` | down/warning counts + worst-CPU nodes in one call |\n\n## SolarWinds writes\n\n| Tool | Risk | Path / verb | Undo / safety |\n|------|------|-------------|---------------|\n| `list_events` | read | `Orion.Events` | recent events (read) |\n| `list_unmanaged` | read | `Orion.Nodes` where Unmanaged | currently-unmanaged nodes (read) |\n| `list_muted` | read | `Orion.AlertSuppression` | currently-muted objects (read) |\n| `mute_alerts` | write **med** | `AlertSuppression` (SuppressAlerts) | **time-boxed** (requires end time); records inverse **unmute** undo |\n| `unmute_alerts` | write **med** | `AlertSuppression` (ResumeAlerts) | un-suppresses alerting |\n| `schedule_maintenance` | write **med** | `AlertSuppression` window | **requires an end time** (time-boxed maintenance window) |\n| `unmanage_node` | write **HIGH** | `Orion.Nodes.Unmanage` verb | `dry_run` + double-confirm; captures prior managed state; records inverse **remanage** undo |\n| `remanage_node` | write **med** | `Orion.Nodes.Remanage` verb | brings a node back under management |\n| `remove_node` | write **HIGH** | SWIS `DELETE` on the node URI | `dry_run` + double-confirm; no undo (deletion is not reversible) |\n\n## PRTG (read)\n\n| Tool | Path | Returns |\n|------|------|---------|\n| `prtg_sensors` | `/api/table.json?content=sensors` | sensors (status, last value, message) |\n| `prtg_sensor_details` | `/api/getsensordetails.json` | one sensor's detail (channels, uptime, last check) |\n| `prtg_devices` | `/api/table.json?content=devices` | devices (host, group, status) |\n| `prtg_groups` | `/api/table.json?content=groups` | probe/group tree with status rollups |\n| `prtg_history` | `/api/historicdata.json` | historic values for a sensor over a window |\n| `prtg_system_status` | `/api/status.json` | server/system status (also the PRTG `doctor` check) |\n| `prtg_alarms` | `/api/table.json?content=messages` (alarms) | active PRTG alarms |\n\n## PRTG writes\n\n| Tool | Risk | Path | Undo / safety |\n|------|------|------|---------------|\n| `pause_sensor` | write **med** | `/api/pause.htm?action=0` | records inverse **resume** undo |\n| `resume_sensor` | write **med** | `/api/pause.htm?action=1` | resumes a paused sensor |\n| `schedule_maintenance_prtg` | write **med** | `/api/pauseobjectfor.htm?duration=` | **time-boxed** (requires minutes) |\n\n## Zabbix (read)\n\n| Tool | JSON-RPC method | Returns |\n|------|-----------------|---------|\n| `zabbix_problems` | `problem.get` | current problems; severity 0-5 mapped to names + canonical levels (`info`/`warning`/`high`/`critical`) |\n| `zabbix_hosts` | `host.get` (+interfaces, +host groups) | host inventory: monitored flag, interfaces (ip/dns/availability), groups |\n| `zabbix_hostgroups` | `hostgroup.get` | host groups (ids + names) |\n| `zabbix_triggers` | `trigger.get` | triggers, by default only those currently firing (PROBLEM) |\n| `zabbix_events` | `event.get` | recent trigger events (newest first, capped) |\n| `zabbix_item_history` | `item.get` + `history.get` | **bounded** metric detail: item meta + history points (window ≤ 168 h, ≤ 500 points) |\n| `zabbix_maintenances` | `maintenance.get` | maintenance windows with hosts/groups + periods |\n\n## Zabbix writes\n\n| Tool | Risk | JSON-RPC method | Undo / safety |\n|------|------|-----------------|---------------|\n| `zabbix_create_maintenance` | write **med** | `maintenance.create` | **time-boxed** (minutes > 0) + must name hosts/groups; records a **replayable undo** = delete exactly the created maintenance id |\n| `zabbix_delete_maintenance` | write **HIGH** | `maintenance.delete` | `dry_run` + double-confirm; the window's **full definition** is captured into priorState first (no undo — re-create manually from it) |\n\n## Out of scope (by design)\n\n- Monitoring stacks other than SolarWinds Orion, PRTG, and Zabbix\n- Zabbix template / discovery / user CRUD, `trend.get`, and host **onboarding**\n- Creating alerts / thresholds / SWQL-view CRUD, and node/interface **onboarding**\n- Anything outside monitoring (hypervisor, storage, backup, cluster, network\n  device config, OT/industrial) — route to the appropriate other AIops-tools skill\n\nWant one of these? Open an issue or PR — feedback and contributions welcome.\n\nFile v0.10.3:references/cli-reference.md\n\n# monitoring-aiops CLI reference\n\n> Covers SolarWinds Orion (SWIS REST + SWQL), Paessler\n> PRTG (web API), and Zabbix 6.x/7.x (JSON-RPC); SWIS/PRTG/Zabbix responses are\n> mocked and need live verification.\n> The CLI is a convenience subset — the full 42-tool surface is via the MCP\n> server (`monitoring-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nmonitoring-aiops init                      # interactive wizard (asks for the platform: solarwinds/prtg/zabbix)\nmonitoring-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   SolarWinds: a SWQL query · PRTG: /api/status.json\n                                           #   Zabbix: apiinfo.version (no auth) + authed host count\nmonitoring-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.monitoring-aiops/secrets.enc)\n\n```bash\nmonitoring-aiops secret set <target> [--value <secret>]  # store Orion password / PRTG or Zabbix token (hidden prompt if no --value)\nmonitoring-aiops secret list                             # names only — secrets never shown\nmonitoring-aiops secret rm <target>\nmonitoring-aiops secret migrate                          # import legacy plaintext env (MONITORING_<TARGET>_SECRET)\nmonitoring-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nmonitoring-aiops overview [--target <t>]   # NOC summary: platform + active/unacked alert counts + top rollup\n```\n\n## SWQL (SolarWinds)\n\n```bash\nmonitoring-aiops swql library                    # list the canned queries\nmonitoring-aiops swql canned <name>              # run a canned query: nodes_down, flapping_interfaces,\n                                                 #   muted_report, high_cpu_nodes, volumes_full, unmanaged_scheduled\nmonitoring-aiops swql query \"SELECT ...\"         # validated read-only SWQL passthrough (SELECT only)\n```\n\n## Alerts (all platforms)\n\n```bash\nmonitoring-aiops alert list [--target <t>]       # active alerts, deduped/rolled up by message\nmonitoring-aiops alert ack <alert_id>            # acknowledge an alert / PRTG alarm / Zabbix problem event\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `swql`, and `alert` are the CLI subset; the remaining SolarWinds\n  health, PRTG, Zabbix, and governed-write tools (mute/unmute,\n  schedule_maintenance, unmanage/remanage/remove node, PRTG pause/resume,\n  Zabbix maintenance create/delete) are exposed through the MCP server.\n  High-risk MCP writes use dry-run + double-confirm; `MONITORING_AUDIT_APPROVED_BY`\n  / `MONITORING_AUDIT_RATIONALE` are recorded on the audit row when set.\n\nFile v0.10.3:references/setup-guide.md\n\n# monitoring-aiops setup & security guide\n\n> Not yet validated against a live NOC (see `docs/VERIFICATION.md`). **PRTG's free\n> perpetual 100-sensor Freeware edition (with the API) and an open-source\n> Zabbix appliance (Docker compose) are the easiest live checks; SolarWinds is\n> a 30-day trial only — mock-only past that, the largest verification debt.**\n\n## 1. Install\n\n```bash\nuv tool install monitoring-aiops\n```\n\n## 2. Get a credential\n\n- **SolarWinds Orion** — an Orion account (username + password) with the Orion\n  API enabled. monitoring-aiops talks to SWIS REST + SWQL with **HTTP Basic\n  auth** on port **17774** — the SWIS port since Orion 2023.1, which deprecated\n  the old **17778** and slated it for removal. Leave `port` unset and a\n  pre-2023.1 server is still reached: the first connection failure triggers one\n  retry on 17778, and the port that answered is reused for that session. Set\n  `port:` yourself and it is used verbatim, with no fallback probing.\n- **PRTG** — an **API token** (Setup → Account Settings → API Keys, or a\n  passhash). PRTG's web API is on port **443/8080**. A free Freeware edition\n  (100 sensors, perpetual) exposes the same API — the easiest way to self-test.\n- **Zabbix (6.x/7.x)** — an **API token** (Administration → API tokens; user\n  tokens under User settings → API tokens). monitoring-aiops talks JSON-RPC 2.0\n  to `/api_jsonrpc.php` on the frontend port (default **443**), sending the\n  token as a `Bearer` header (6.4+/7.x) with an automatic legacy `auth`-field\n  fallback for 6.0. Zabbix is fully open source — a Docker-compose appliance\n  is a 10-minute self-test.\n\n## 3. Onboard\n\n```bash\nmonitoring-aiops init\n```\n\nThe wizard asks, per target, for the **platform** (`solarwinds` / `prtg` /\n`zabbix`), the **host**, the **port** (defaults 17774 for SolarWinds, 443 for\nPRTG and Zabbix — accept the default and no `port:` key is written, so the\ndefault keeps tracking the platform), the **Orion username** (SolarWinds only),\nand the **secret**\n— the Orion account password, the PRTG API token, or the Zabbix API token.\nNon-secret connection details go to `~/.monitoring-aiops/config.yaml`; the\nsecret is stored **encrypted** into `~/.monitoring-aiops/secrets.enc`. Example\nconfig (one config can span all NOCs):\n\n```yaml\ntargets:\n  - name: orion1\n    platform: solarwinds\n    host: 10.0.0.20\n    # port omitted -> 17774 (Orion 2023.1+), falling back to 17778 once if\n    # nothing answers. Set it here only to pin a non-standard SWIS port.\n    username: admin\n    verify_ssl: false          # self-signed lab certs only\n  - name: prtg1\n    platform: prtg\n    host: 10.0.0.40\n    port: 443\n    verify_ssl: true\n  - name: zbx1\n    platform: zabbix\n    host: 10.0.0.60\n    port: 443\n    verify_ssl: true\n```\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nExport the master password so the encrypted store can be unlocked without a\nprompt:\n\n```bash\nexport MONITORING_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## Credential security\n\n- The secret (Orion password / PRTG API token / Zabbix API token) is **never**\n  written to disk in\n  plaintext. It lives only in `~/.monitoring-aiops/secrets.enc`, encrypted with\n  Fernet (AES-128-CBC + HMAC), the key derived from your master password via\n  scrypt. Only a per-store random salt and the ciphertext are on disk (chmod\n  600); the master password itself is never stored.\n- A legacy plaintext env var `MONITORING_<TARGET_NAME_UPPER>_SECRET` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `monitoring-aiops secret migrate` (it imports then renames the old `.env`).\n- The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG / Zabbix\n  API token at request time and held only in memory; it is never logged or\n  echoed. Exception\n  text and tracebacks are scrubbed of secret-shaped strings before being written\n  to the audit log.\n\n## Governance harness state\n\nState lives under `~/.monitoring-aiops/` (relocate with `MONITORING_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with risk tier, approver, rationale\n- `undo.db` — inverse descriptors for reversible writes (mute→unmute,\n  unmanage→remanage, pause→resume, zabbix create-maintenance→delete)\n- budget / runaway guard — caps cumulative tool calls and wall-time; trips on\n  tight poll/retry loops\n\n## Governed writes\n\n- **High-risk** ops (`unmanage_node`, `remove_node`,\n  `zabbix_delete_maintenance`) use `dry_run` + double confirmation;\n  `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional\n  audit annotations, recorded when set but never required.\n- **Time-boxed** ops require an end time / duration: `mute_alerts`,\n  `schedule_maintenance` (SolarWinds), `schedule_maintenance_prtg` (PRTG, in\n  minutes), and `zabbix_create_maintenance` (Zabbix, in minutes). This prevents\n  forgotten, indefinite suppression windows.\n\n## Verify\n\n```bash\nmonitoring-aiops doctor\n```\n\n`doctor` is platform-aware: it checks the config file, the encrypted store and\nits permissions, that a secret is present per target, and (unless `--skip-auth`)\nconnectivity — a SWQL query for SolarWinds targets, `/api/status.json` for PRTG\ntargets, and for Zabbix targets the unauthenticated `apiinfo.version`\n(reachability) followed by a cheap authed host count (token validity).\n\nFile v0.10.3:skill-card.md\n\n## Description:\n\nMonitoring AIops helps agents operate SolarWinds Orion, PRTG, and Zabbix monitoring environments with NOC summaries, alert rollups, health checks, read-only SWQL queries, and audited maintenance or suppression actions.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, SREs, and NOC operators use this skill to inspect and operate SolarWinds Orion, PRTG, and Zabbix monitoring estates, including alert triage, health checks, read-only SWQL questions, and guarded maintenance workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can perform high-impact monitoring changes, including acknowledgement, suppression, maintenance, unmanage, pause, resume, and deletion workflows.\n\nMitigation: Start with least-privilege, preferably read-only monitoring credentials, and grant write-capable credentials only when the agent runtime has separate approval and change-control safeguards.\n\nRisk: Production targets may be exposed to weak transport controls if lab-friendly TLS settings are carried forward.\n\nMitigation: Enable TLS verification for production SolarWinds, PRTG, and Zabbix targets.\n\nRisk: The master password can be supplied through an environment variable for non-interactive use.\n\nMitigation: Avoid broad process environments for master passwords; prefer tightly scoped runtime secrets or an interactive prompt where practical.\n\nRisk: Artifact documentation says behavior is exercised against mocked SWIS, PRTG, and Zabbix responses and still needs live NOC verification.\n\nMitigation: Run live checks in a test monitoring environment before production use, starting with read-only workflows and the documented doctor command.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/monitoring-aiops)\n- [Monitoring AIops source repository](https://github.com/AIops-tools/Monitoring-AIops)\n- [Capabilities reference](references/capabilities.md)\n- [Setup and security guide](references/setup-guide.md)\n- [CLI reference](references/cli-reference.md)\n- [Agent guardrails](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Shell commands, Configuration, Guidance]\n\n**Output Format:** [Markdown or plain text with monitoring observations, CLI commands, configuration snippets, and structured tool-result summaries.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include platform identifiers, status values, limits, truncation flags, and audit or undo references.]\n\n## Skill Version(s):\n\n0.10.3 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.2: 7 files, 19394 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2882b), SKILL.md (16987b), _meta.json (136b)\n\nFile v0.10.2:SKILL.md\n\n---\nname: monitoring-aiops\nslug: monitoring-aiops\ndisplayName: \"Monitoring AIops\"\nsummary: \"Governed SolarWinds Orion + PRTG + Zabbix ops: SWQL, alert rollup, health, 42 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Monitoring-AIops\ntags: [aiops, mcp, governance, monitoring]\ndescription: >\n  Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window).\n  Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring.\n  Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill.\n  Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.\ninstaller:\n  kind: uv\n  package: monitoring-aiops\nargument-hint: \"[node/sensor id, a SWQL question, or describe your NOC task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"monitoring-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"MONITORING_AIOPS_CONFIG\",\"MONITORING_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Monitoring-AIops\",\"emoji\":\"📡\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed monitoring operations across SolarWinds Orion (SWIS REST + SWQL, port 17774 on Orion 2023.1+ with an automatic one-shot fallback to the legacy 17778, HTTP Basic auth), Paessler PRTG (web API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at /api_jsonrpc.php, API token as Bearer header on 6.4+/7.x with a legacy auth-field fallback for 6.0). Each target in the config names its own platform, so one config can span all NOCs. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.monitoring-aiops/ (relocatable via MONITORING_AIOPS_HOME).\n  Credentials: the Orion account password (SolarWinds), the PRTG API token, or the Zabbix API token is stored ENCRYPTED in ~/.monitoring-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'monitoring-aiops init' to onboard (it asks for the platform), or 'monitoring-aiops secret set <target>' to add one. The store is unlocked by a master password from MONITORING_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var MONITORING_<TARGET_NAME_UPPER>_SECRET is still honoured as a fallback with a deprecation warning (migrate with 'monitoring-aiops secret migrate'). The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG/Zabbix API token at request time and held only in memory; secrets are never logged or echoed.\n  Read-only SWQL passthrough (swql_query) is validated to accept SELECT statements only. State-changing operations pass through the @governed_tool decorator (budget guard + audit + undo recording; each tool's risk_level is recorded as a descriptive tier, not a gate). Destructive writes (unmanage_node, remove_node, zabbix_delete_maintenance) are high-risk with dry_run + double confirmation; unmanage_node records an inverse remanage undo descriptor, zabbix_delete_maintenance captures the window's FULL definition into priorState first. Suppression/maintenance writes are TIME-BOXED (mute_alerts, schedule_maintenance, schedule_maintenance_prtg, zabbix_create_maintenance require an end time / duration). mute_alerts→unmute, pause_sensor→resume, zabbix_create_maintenance→delete-that-maintenance-id record inverse undo descriptors. Zabbix item history is BOUNDED (capped window + point count).\n  Webhooks: none — no outbound network calls beyond the configured SolarWinds SWIS / PRTG web API / Zabbix JSON-RPC endpoint.\n  SSL: verify_ssl defaults to false-friendly for self-signed lab certs; enable for production.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked SWIS/PRTG/Zabbix responses; not yet run against a live NOC (see docs/VERIFICATION.md). PRTG has a free perpetual 100-sensor Freeware edition with the API, and Zabbix is fully open source (a Docker-compose appliance is a 10-minute live check) — the easiest live checks; SolarWinds is a 30-day trial (mock-only past that — largest verification debt).\n---\n\n# Monitoring AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by SolarWinds, Paessler, Zabbix, or any monitoring vendor.** SolarWinds, Orion, SWQL, THWACK, PRTG, Paessler and Zabbix are trademarks of their respective owners. Source at [github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops) under the MIT license.\n\nGoverned network / infrastructure monitoring operations — **42 MCP tools**\nacross **SolarWinds Orion** (SWIS REST + SWQL), **Paessler PRTG** (web API),\nand **Zabbix 6.x/7.x** (JSON-RPC 2.0),\nevery one wrapped with the bundled `@governed_tool` harness: a local unified\naudit log under `~/.monitoring-aiops/`, policy engine, token/runaway budget\nguard, undo-token recording, and risk-tier labelling on the audit trail. One\nconfig can span all NOCs. The Orion password / PRTG API token / Zabbix API token is stored\n**encrypted** (`~/.monitoring-aiops/secrets.enc`, Fernet + scrypt) — never\nplaintext on disk.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`monitoring_aiops.governance`) — no external skill-family dependency.\n> PRTG's free Freeware edition and an open-source Zabbix appliance are the\n> easiest live checks; SolarWinds is trial-only past 30 days (largest\n> verification debt — see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **SWQL** | SolarWinds | library, canned, query (SELECT-only passthrough) | 3 | read |\n| **Alerts** | all | active_alerts (dedup/rollup), alert_acknowledge | 2 | 1 read, 1 write |\n| **SolarWinds health** | SolarWinds | node/nodes/interface/volume/application status, topn, noc_rollup | 7 | read |\n| **SolarWinds writes** | SolarWinds | list_events/unmanaged/muted | 3 | read |\n| | SolarWinds | mute/unmute, schedule_maintenance, remanage_node | 4 | write (med) |\n| | SolarWinds | unmanage_node, remove_node | 2 | write (**high**) |\n| **PRTG** | PRTG | sensors/sensor_details/devices/groups/history/system_status/alarms | 7 | read |\n| **PRTG writes** | PRTG | pause_sensor, resume_sensor, schedule_maintenance_prtg | 3 | write (med) |\n| **Zabbix** | Zabbix | zabbix_problems/hosts/hostgroups/triggers/events/item_history/maintenances | 7 | read |\n| **Zabbix writes** | Zabbix | zabbix_create_maintenance (time-boxed; undo = delete that id) | 1 | write (med) |\n| | Zabbix | zabbix_delete_maintenance (priorState = full definition) | 1 | write (**high**) |\n| **Undo** | all | undo_list, undo_apply | 2 | undo |\n\nThe canned SWQL library (`swql_library` lists them) answers the most-repeated\nTHWACK questions directly: `nodes_down`, `flapping_interfaces`, `muted_report`,\n`high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled`. For anything else,\n`swql_query` is a validated read-only (SELECT-only) SWQL passthrough.\n\n## Quick Install\n\n```bash\nuv tool install monitoring-aiops\nmonitoring-aiops init       # wizard: pick platform (solarwinds/prtg/zabbix) + encrypted secret\nmonitoring-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/monitoring-aiops\nopenclaw skills info monitoring-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a NOC snapshot (`overview` / `noc_rollup`): active/unacked alert counts,\n  down/warning nodes, worst CPU\n- Answer a repeated SWQL question (`swql_library` → `swql_canned nodes_down`), or\n  run an ad-hoc read-only SWQL SELECT (`swql_query`)\n- Triage an alert storm (`active_alerts` dedup/rollup collapses flap/down\n  storms), then `alert_acknowledge`\n- SolarWinds health: `node_status`, `interface_status` (top-N by util),\n  `volume_status`, `application_status` (SAM), `topn` (cpu/mem/latency/loss)\n- PRTG: list `prtg_sensors` / `prtg_devices` / `prtg_groups`, drill with\n  `prtg_sensor_details` / `prtg_history`, check `prtg_alarms` / `prtg_system_status`\n- Zabbix: triage `zabbix_problems` (0-5 severity mapped to levels) /\n  `zabbix_triggers`, inventory `zabbix_hosts` / `zabbix_hostgroups`, drill with\n  `zabbix_item_history` (bounded), review `zabbix_events` / `zabbix_maintenances`\n- Safely take a node out for maintenance (`schedule_maintenance` /\n  `unmanage_node` with dry_run + double-confirm), pause a PRTG sensor\n  (`pause_sensor`), or create a time-boxed Zabbix maintenance window\n  (`zabbix_create_maintenance` — undo deletes exactly that window)\n\n**Do NOT use when** the target is not a SolarWinds/PRTG/Zabbix monitoring\nplatform — route hypervisor, storage, backup, cluster, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| SolarWinds Orion / SWQL, PRTG, or Zabbix monitoring ops | **monitoring-aiops** (this skill) |\n| A non-monitoring platform (hypervisor, storage, backup, cluster, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Other monitoring stacks (not SolarWinds/PRTG/Zabbix) | out of scope for this tool |\n\n## Common Workflows\n\n> **No authorization gate**: the skill runs the operations you ask for and audits every one; it does not decide whether a write is permitted — that is the agent's judgement or the permissions of the SolarWinds/PRTG/Zabbix account it connects with (a read-only monitoring account makes writes fail at the server). There is no read-only switch, policy file, or approval gate. `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n\n### 1. The 3 a.m. alert storm — collapse it, then acknowledge what matters\n\n1. `monitoring-aiops doctor` → confirm the NOC platform is actually reachable\n   (a \"storm\" is sometimes just a poller that lost the target)\n2. `monitoring-aiops overview` → the one-screen picture: down/warning counts\n   across the configured targets\n3. `monitoring-aiops alert list` (MCP: `active_alerts`) → deduped / rolled-up\n   entries; an interface-flap or node-down storm collapses into **one** entry\n   with a count instead of a wall of alerts\n4. `noc_rollup` → confirm whether the storm has a single upstream cause (one\n   node down taking its children with it) rather than N independent faults\n5. Acknowledge only the rolled-up entry that matters:\n   `monitoring-aiops alert ack <alert-id>` (SolarWinds `AlertActive.Acknowledge`\n   / PRTG `acknowledgealarm` / Zabbix `event.acknowledge`) — the prior ack state\n   is captured into priorState, and the ack is double-confirmed\n6. **Failure branch**: if `doctor` fails, do **not** acknowledge anything — you\n   would be silencing alerts you cannot currently see. Fix credentials with\n   `monitoring-aiops secret set <target>` first. If you acknowledged the wrong\n   alert, `monitoring-aiops undo list` → `undo apply <id>` restores the prior\n   ack state.\n\n### 2. \"Which nodes are down and what's saturated?\" (read-only)\n\n1. `noc_rollup` → down / warning counts plus the worst-CPU nodes in a single\n   call, so you do not page through a dashboard\n2. `topn cpu` (also `memory`, `latency`, `packetloss`) → the worst offenders\n   with the measured number\n3. `node_status <node>` → drill into one node; `interface_status` for a\n   suspected link problem, `volume_status` for a filling disk,\n   `application_status` for an app-layer fault\n4. `list_events` → what changed around the time things went bad\n5. `list_unmanaged` → check whether a \"missing\" node is simply unmanaged from a\n   previous maintenance window that was never reverted\n6. **Failure branch**: if a node shows down but is reachable from your shell,\n   the fault is in polling, not the node — check `list_muted` and\n   `list_unmanaged` before escalating to the network team.\n\n### 3. Planned maintenance: suppress noise time-boxed, then restore\n\n1. `node_status <node>` / `swql_canned nodes_down` → confirm you have the right\n   node and that it is currently healthy (so you can tell the difference\n   afterwards)\n2. Prefer the **time-boxed** path — it expires on its own:\n   `schedule_maintenance <node> --end ...` (SolarWinds),\n   `schedule_maintenance_prtg` (PRTG), or `zabbix_create_maintenance` (Zabbix,\n   undo → delete that maintenance id)\n3. If you genuinely need to unmanage instead:\n   `unmanage_node <node> --dry-run`, then re-run without `--dry-run` →\n   **high** risk, double confirmation; it records an inverse `remanage_node`\n   undo descriptor\n4. For a single noisy sensor rather than a whole node: `pause_sensor` (PRTG,\n   undo → `resume_sensor`) or `mute_alerts` (undo → `unmute_alerts`)\n5. When maintenance ends: `remanage_node <node>` / `resume_sensor` /\n   `unmute_alerts`, or simply `monitoring-aiops undo apply <id>` to replay the\n   recorded inverse\n6. **Failure branch**: the classic failure here is *forgetting to restore* —\n   run `list_unmanaged` and `list_muted` at the end of every maintenance window;\n   anything still listed is silently unmonitored. Time-boxed maintenance windows\n   are preferred precisely because they fail safe.\n\n### 4. Answer a bespoke NOC question with SWQL\n\n1. `monitoring-aiops swql library` (MCP: `swql_library`) → the canned queries,\n   so you do not hand-write what already exists\n2. `monitoring-aiops swql canned nodes_down` → run a canned one directly\n   (also `high_cpu_nodes` and the rest of the library)\n3. Not canned? `monitoring-aiops swql query \"SELECT ...\"` → the passthrough\n   **validates the statement is a read-only SELECT** before it runs; anything\n   else is refused\n4. Feed the result into an action — e.g. a node the query surfaced goes into\n   workflow 3 for a maintenance window\n5. **Failure branch**: a rejected query is almost always a non-SELECT statement\n   or a SWQL/SQL dialect slip (SWQL has no `*` expansion on some entities).\n   Start from the nearest canned query in `swql library` and modify it rather\n   than writing from scratch. The passthrough will not be talked into a write —\n   writes go through the governed tools, where they are audited.\n\n## Governance & Safety\n\n- Every tool is audited to `~/.monitoring-aiops/audit.db` (relocatable via\n  `MONITORING_AIOPS_HOME`).\n- Each tool's `risk_level` is carried into the audit row as a descriptive tier\n  (a label, not a gate). `MONITORING_AUDIT_APPROVED_BY` /\n  `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n- Destructive writes support `--dry-run` and double confirmation at the CLI.\n- Suppression / maintenance writes are **time-boxed** (require an end time /\n  duration). Reversible writes record an inverse descriptor (mute→unmute,\n  unmanage→remanage, pause→resume, zabbix_create_maintenance→delete that\n  maintenance id). `zabbix_delete_maintenance` captures the window's full\n  definition into priorState before deleting.\n\n## References\n\n- `references/capabilities.md` — full tool + platform + SWQL/API-path reference\n- `references/cli-reference.md` — CLI command reference\n- `references/setup-guide.md` — onboarding, credentials, and connectivity\n\nFile v0.10.2:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"monitoring-aiops\",\n  \"version\": \"0.10.2\",\n  \"publishedAt\": 1789223383973\n}\n\nFile v0.10.2:references/agent-guardrails.md\n\n# Agent guardrails — running monitoring-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The account you connect with.** Give it a SolarWinds/PRTG/Zabbix login with\n  read-only monitoring scope. A write then fails at the server, which is the\n  only place the permission actually lives — no skill-side flag can be argued\n  around by a model, but a revoked permission cannot be.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Never write to the monitoring database\" | `swql_query` accepts a single read-only `SELECT` and nothing else — no verb invoke, no multi-statement, no DELETE. Orion state changes only happen through the named, governed write tools. |\n| \"Don't invent a value when a field is missing\" | A column the platform did not return comes back as `null`, never as `\"\"`. An absent Orion `StatusDescription`, a PRTG sensor `message`, a Zabbix host `dns` — all distinguishable from a genuinely empty one. |\n| \"Tell me if the output was cut off\" | Every row-capped read returns `returned` / `limit` / `truncated`: `swql_query`, `swql_canned`, `list_events`, `zabbix_events`, `zabbix_item_history`, and `interface_status` with a `top`. Truncation is measured — one row past the cap is fetched, or the full set is counted before the cut — never guessed from a length coincidence. |\n| \"Deduplicate the alert storm before showing me\" | `active_alerts` already rolls repeats of the same message into one row with a `count` and up to three `examples`, worst-first. Report the rollup; do not re-count the raw list. |\n| \"Normalise severity across platforms\" | Zabbix's 0–5 scale is already mapped to canonical `level` values (`info`/`warning`/`high`/`critical`) alongside the platform's own `severity` name. Use `level` for cross-platform statements and `severity` when quoting the platform. |\n| \"Confirm before anything disruptive\" | `remove_node`, `unmanage_node`, `mute_alerts`, and the maintenance-window writes require a `--dry-run`-able preview + double confirmation at the CLI. |\n| \"Log what you did\" | Every call is audited to `~/.monitoring-aiops/audit.db` regardless of what the model says it did. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate an enterprise monitoring system through the monitoring-aiops MCP\ntools. A target is SolarWinds Orion, Paessler PRTG, or Zabbix — check which\nbefore reasoning about what a field means.\n\nTOOL USE\n- Before answering any question about the current monitored estate, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"everything is fine\".\n- Prefer the named canned SWQL (swql_library / swql_canned) over writing your\n  own query. Hand-written SWQL is where small models most often produce\n  syntactically valid nonsense.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete. An alert count from a truncated read is not\n  the number of alerts.\n- A null field means the platform did not return that value. Report it as \"not\n  available\" — never infer it.\n- Report values exactly as returned. Do not translate status codes, severity\n  names, node captions, or sensor names into your own vocabulary.\n- Acknowledged is not the same as resolved. An acknowledged alert is still\n  active; say which you mean.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, root cause, or business impact unless a tool result\n  supports it. A down sensor is one sensor, not necessarily a down service.\n- Do not confuse the identifier kinds: an Orion node Caption, an AlertActiveID,\n  a PRTG objid, and a Zabbix eventid/triggerid/itemid are different namespaces\n  and are not interchangeable across targets.\n- Muting, unmanaging, and maintenance windows suppress alerting; they do not fix\n  anything. Never describe them as a resolution.\n```\n\n## Recommended setup for a local model\n\nStart with a connection that *cannot* write — a SolarWinds/PRTG/Zabbix account\nwith read-only monitoring scope — verify, and widen the account's permission\nonly when you trust the setup:\n\n```bash\nmonitoring-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport MONITORING_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport MONITORING_AUDIT_RATIONALE=\"change window CHG0041231, muting core-sw1\"\n```\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Lead with `noc_rollup` (SolarWinds)\n  or `active_alerts` — they do the correlation and dedup inside one call, so the\n  model does not have to chain reads and keep ids straight.\n- **The model writes broken SWQL.** Use `swql_library` and `swql_canned` — the\n  canned queries answer the most-asked questions and are already parameterised.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions and use `top` / `limit` deliberately rather than pulling every\n  sensor in the estate.\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.10.2:references/capabilities.md\n\n# monitoring-aiops capabilities\n\n> **42 MCP tools** (30 read, 10 write, 2 undo) across SolarWinds\n> Orion (SWIS REST + SWQL, port 17774 with a legacy-17778 fallback, HTTP\n> Basic auth), Paessler PRTG (web\n> API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at\n> `/api_jsonrpc.php`, API token — Bearer header on 6.4+/7.x, legacy `auth`\n> field fallback for 6.0). Each config target names its own `platform`.\n> SWIS/PRTG/Zabbix responses are mocked and need live verification.\n\n## SWQL — SolarWinds (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `swql_library` | (local) | the catalogue of canned queries: `nodes_down`, `flapping_interfaces`, `muted_report`, `high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled` |\n| `swql_canned` | named SWQL → SWIS `/Query` | rows for the named canned query |\n| `swql_query` | validated read-only SWQL → SWIS `/Query` | rows for a caller SELECT (SELECT-only; rejected otherwise) |\n\n## Alerts — all platforms\n\n| Tool | Risk | Path | Returns / effect |\n|------|------|------|------------------|\n| `active_alerts` | read | SWIS `AlertActive`/`AlertObjects`, PRTG `/api/table.json?content=messages`, or Zabbix `problem.get` | active alerts **deduped/rolled up by message** — flap/down storms collapse into one counted entry |\n| `alert_acknowledge` | write **medium** | SW `AlertActive.Acknowledge` verb / PRTG `acknowledgealarm.htm` / Zabbix `event.acknowledge` (action 6; prior ack state → priorState) | acknowledges an alert / alarm / problem event |\n\n## SolarWinds health (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `node_status` | `Orion.Nodes` | one node's status, CPU/mem, response time |\n| `nodes_list` | `Orion.Nodes` | node inventory (status, vendor, IP, last boot) |\n| `interface_status` | `Orion.NPM.Interfaces` | top-N interfaces by utilisation (in/out, errors, oper status) |\n| `volume_status` | `Orion.Volumes` | volumes by % used (size, used, type) |\n| `application_status` | `Orion.APM.Application` (SAM) | SAM application/component status |\n| `topn` | `Orion.Nodes` metrics | top-N nodes by `cpu` / `memory` / `latency` / `packetloss` |\n| `noc_rollup` | folds `Orion.Nodes` | down/warning counts + worst-CPU nodes in one call |\n\n## SolarWinds writes\n\n| Tool | Risk | Path / verb | Undo / safety |\n|------|------|-------------|---------------|\n| `list_events` | read | `Orion.Events` | recent events (read) |\n| `list_unmanaged` | read | `Orion.Nodes` where Unmanaged | currently-unmanaged nodes (read) |\n| `list_muted` | read | `Orion.AlertSuppression` | currently-muted objects (read) |\n| `mute_alerts` | write **med** | `AlertSuppression` (SuppressAlerts) | **time-boxed** (requires end time); records inverse **unmute** undo |\n| `unmute_alerts` | write **med** | `AlertSuppression` (ResumeAlerts) | un-suppresses alerting |\n| `schedule_maintenance` | write **med** | `AlertSuppression` window | **requires an end time** (time-boxed maintenance window) |\n| `unmanage_node` | write **HIGH** | `Orion.Nodes.Unmanage` verb | `dry_run` + double-confirm; captures prior managed state; records inverse **remanage** undo |\n| `remanage_node` | write **med** | `Orion.Nodes.Remanage` verb | brings a node back under management |\n| `remove_node` | write **HIGH** | SWIS `DELETE` on the node URI | `dry_run` + double-confirm; no undo (deletion is not reversible) |\n\n## PRTG (read)\n\n| Tool | Path | Returns |\n|------|------|---------|\n| `prtg_sensors` | `/api/table.json?content=sensors` | sensors (status, last value, message) |\n| `prtg_sensor_details` | `/api/getsensordetails.json` | one sensor's detail (channels, uptime, last check) |\n| `prtg_devices` | `/api/table.json?content=devices` | devices (host, group, status) |\n| `prtg_groups` | `/api/table.json?content=groups` | probe/group tree with status rollups |\n| `prtg_history` | `/api/historicdata.json` | historic values for a sensor over a window |\n| `prtg_system_status` | `/api/status.json` | server/system status (also the PRTG `doctor` check) |\n| `prtg_alarms` | `/api/table.json?content=messages` (alarms) | active PRTG alarms |\n\n## PRTG writes\n\n| Tool | Risk | Path | Undo / safety |\n|------|------|------|---------------|\n| `pause_sensor` | write **med** | `/api/pause.htm?action=0` | records inverse **resume** undo |\n| `resume_sensor` | write **med** | `/api/pause.htm?action=1` | resumes a paused sensor |\n| `schedule_maintenance_prtg` | write **med** | `/api/pauseobjectfor.htm?duration=` | **time-boxed** (requires minutes) |\n\n## Zabbix (read)\n\n| Tool | JSON-RPC method | Returns |\n|------|-----------------|---------|\n| `zabbix_problems` | `problem.get` | current problems; severity 0-5 mapped to names + canonical levels (`info`/`warning`/`high`/`critical`) |\n| `zabbix_hosts` | `host.get` (+interfaces, +host groups) | host inventory: monitored flag, interfaces (ip/dns/availability), groups |\n| `zabbix_hostgroups` | `hostgroup.get` | host groups (ids + names) |\n| `zabbix_triggers` | `trigger.get` | triggers, by default only those currently firing (PROBLEM) |\n| `zabbix_events` | `event.get` | recent trigger events (newest first, capped) |\n| `zabbix_item_history` | `item.get` + `history.get` | **bounded** metric detail: item meta + history points (window ≤ 168 h, ≤ 500 points) |\n| `zabbix_maintenances` | `maintenance.get` | maintenance windows with hosts/groups + periods |\n\n## Zabbix writes\n\n| Tool | Risk | JSON-RPC method | Undo / safety |\n|------|------|-----------------|---------------|\n| `zabbix_create_maintenance` | write **med** | `maintenance.create` | **time-boxed** (minutes > 0) + must name hosts/groups; records a **replayable undo** = delete exactly the created maintenance id |\n| `zabbix_delete_maintenance` | write **HIGH** | `maintenance.delete` | `dry_run` + double-confirm; the window's **full definition** is captured into priorState first (no undo — re-create manually from it) |\n\n## Out of scope (by design)\n\n- Monitoring stacks other than SolarWinds Orion, PRTG, and Zabbix\n- Zabbix template / discovery / user CRUD, `trend.get`, and host **onboarding**\n- Creating alerts / thresholds / SWQL-view CRUD, and node/interface **onboarding**\n- Anything outside monitoring (hypervisor, storage, backup, cluster, network\n  device config, OT/industrial) — route to the appropriate other AIops-tools skill\n\nWant one of these? Open an issue or PR — feedback and contributions welcome.\n\nFile v0.10.2:references/cli-reference.md\n\n# monitoring-aiops CLI reference\n\n> Covers SolarWinds Orion (SWIS REST + SWQL), Paessler\n> PRTG (web API), and Zabbix 6.x/7.x (JSON-RPC); SWIS/PRTG/Zabbix responses are\n> mocked and need live verification.\n> The CLI is a convenience subset — the full 42-tool surface is via the MCP\n> server (`monitoring-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nmonitoring-aiops init                      # interactive wizard (asks for the platform: solarwinds/prtg/zabbix)\nmonitoring-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   SolarWinds: a SWQL query · PRTG: /api/status.json\n                                           #   Zabbix: apiinfo.version (no auth) + authed host count\nmonitoring-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.monitoring-aiops/secrets.enc)\n\n```bash\nmonitoring-aiops secret set <target> [--value <secret>]  # store Orion password / PRTG or Zabbix token (hidden prompt if no --value)\nmonitoring-aiops secret list                             # names only — secrets never shown\nmonitoring-aiops secret rm <target>\nmonitoring-aiops secret migrate                          # import legacy plaintext env (MONITORING_<TARGET>_SECRET)\nmonitoring-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nmonitoring-aiops overview [--target <t>]   # NOC summary: platform + active/unacked alert counts + top rollup\n```\n\n## SWQL (SolarWinds)\n\n```bash\nmonitoring-aiops swql library                    # list the canned queries\nmonitoring-aiops swql canned <name>              # run a canned query: nodes_down, flapping_interfaces,\n                                                 #   muted_report, high_cpu_nodes, volumes_full, unmanaged_scheduled\nmonitoring-aiops swql query \"SELECT ...\"         # validated read-only SWQL passthrough (SELECT only)\n```\n\n## Alerts (all platforms)\n\n```bash\nmonitoring-aiops alert list [--target <t>]       # active alerts, deduped/rolled up by message\nmonitoring-aiops alert ack <alert_id>            # acknowledge an alert / PRTG alarm / Zabbix problem event\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `swql`, and `alert` are the CLI subset; the remaining SolarWinds\n  health, PRTG, Zabbix, and governed-write tools (mute/unmute,\n  schedule_maintenance, unmanage/remanage/remove node, PRTG pause/resume,\n  Zabbix maintenance create/delete) are exposed through the MCP server.\n  High-risk MCP writes use dry-run + double-confirm; `MONITORING_AUDIT_APPROVED_BY`\n  / `MONITORING_AUDIT_RATIONALE` are recorded on the audit row when set.\n\nFile v0.10.2:references/setup-guide.md\n\n# monitoring-aiops setup & security guide\n\n> Not yet validated against a live NOC (see `docs/VERIFICATION.md`). **PRTG's free\n> perpetual 100-sensor Freeware edition (with the API) and an open-source\n> Zabbix appliance (Docker compose) are the easiest live checks; SolarWinds is\n> a 30-day trial only — mock-only past that, the largest verification debt.**\n\n## 1. Install\n\n```bash\nuv tool install monitoring-aiops\n```\n\n## 2. Get a credential\n\n- **SolarWinds Orion** — an Orion account (username + password) with the Orion\n  API enabled. monitoring-aiops talks to SWIS REST + SWQL with **HTTP Basic\n  auth** on port **17774** — the SWIS port since Orion 2023.1, which deprecated\n  the old **17778** and slated it for removal. Leave `port` unset and a\n  pre-2023.1 server is still reached: the first connection failure triggers one\n  retry on 17778, and the port that answered is reused for that session. Set\n  `port:` yourself and it is used verbatim, with no fallback probing.\n- **PRTG** — an **API token** (Setup → Account Settings → API Keys, or a\n  passhash). PRTG's web API is on port **443/8080**. A free Freeware edition\n  (100 sensors, perpetual) exposes the same API — the easiest way to self-test.\n- **Zabbix (6.x/7.x)** — an **API token** (Administration → API tokens; user\n  tokens under User settings → API tokens). monitoring-aiops talks JSON-RPC 2.0\n  to `/api_jsonrpc.php` on the frontend port (default **443**), sending the\n  token as a `Bearer` header (6.4+/7.x) with an automatic legacy `auth`-field\n  fallback for 6.0. Zabbix is fully open source — a Docker-compose appliance\n  is a 10-minute self-test.\n\n## 3. Onboard\n\n```bash\nmonitoring-aiops init\n```\n\nThe wizard asks, per target, for the **platform** (`solarwinds` / `prtg` /\n`zabbix`), the **host**, the **port** (defaults 17774 for SolarWinds, 443 for\nPRTG and Zabbix — accept the default and no `port:` key is written, so the\ndefault keeps tracking the platform), the **Orion username** (SolarWinds only),\nand the **secret**\n— the Orion account password, the PRTG API token, or the Zabbix API token.\nNon-secret connection details go to `~/.monitoring-aiops/config.yaml`; the\nsecret is stored **encrypted** into `~/.monitoring-aiops/secrets.enc`. Example\nconfig (one config can span all NOCs):\n\n```yaml\ntargets:\n  - name: orion1\n    platform: solarwinds\n    host: 10.0.0.20\n    # port omitted -> 17774 (Orion 2023.1+), falling back to 17778 once if\n    # nothing answers. Set it here only to pin a non-standard SWIS port.\n    username: admin\n    verify_ssl: false          # self-signed lab certs only\n  - name: prtg1\n    platform: prtg\n    host: 10.0.0.40\n    port: 443\n    verify_ssl: true\n  - name: zbx1\n    platform: zabbix\n    host: 10.0.0.60\n    port: 443\n    verify_ssl: true\n```\n\n## 4. Non-interactive use (MCP server / CI / cron)\n\nExport the master password so the encrypted store can be unlocked without a\nprompt:\n\n```bash\nexport MONITORING_AIOPS_MASTER_PASSWORD='your-master-password'\n```\n\n## Credential security\n\n- The secret (Orion password / PRTG API token / Zabbix API token) is **never**\n  written to disk in\n  plaintext. It lives only in `~/.monitoring-aiops/secrets.enc`, encrypted with\n  Fernet (AES-128-CBC + HMAC), the key derived from your master password via\n  scrypt. Only a per-store random salt and the ciphertext are on disk (chmod\n  600); the master password itself is never stored.\n- A legacy plaintext env var `MONITORING_<TARGET_NAME_UPPER>_SECRET` is still\n  honoured as a fallback with a deprecation warning — migrate with\n  `monitoring-aiops secret migrate` (it imports then renames the old `.env`).\n- The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG / Zabbix\n  API token at request time and held only in memory; it is never logged or\n  echoed. Exception\n  text and tracebacks are scrubbed of secret-shaped strings before being written\n  to the audit log.\n\n## Governance harness state\n\nState lives under `~/.monitoring-aiops/` (relocate with `MONITORING_AIOPS_HOME`):\n\n- `audit.db` — every tool call (SQLite), with risk tier, approver, rationale\n- `undo.db` — inverse descriptors for reversible writes (mute→unmute,\n  unmanage→remanage, pause→resume, zabbix create-maintenance→delete)\n- budget / runaway guard — caps cumulative tool calls and wall-time; trips on\n  tight poll/retry loops\n\n## Governed writes\n\n- **High-risk** ops (`unmanage_node`, `remove_node`,\n  `zabbix_delete_maintenance`) use `dry_run` + double confirmation;\n  `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional\n  audit annotations, recorded when set but never required.\n- **Time-boxed** ops require an end time / duration: `mute_alerts`,\n  `schedule_maintenance` (SolarWinds), `schedule_maintenance_prtg` (PRTG, in\n  minutes), and `zabbix_create_maintenance` (Zabbix, in minutes). This prevents\n  forgotten, indefinite suppression windows.\n\n## Verify\n\n```bash\nmonitoring-aiops doctor\n```\n\n`doctor` is platform-aware: it checks the config file, the encrypted store and\nits permissions, that a secret is present per target, and (unless `--skip-auth`)\nconnectivity — a SWQL query for SolarWinds targets, `/api/status.json` for PRTG\ntargets, and for Zabbix targets the unauthenticated `apiinfo.version`\n(reachability) followed by a cheap authed host count (token validity).\n\nFile v0.10.2:skill-card.md\n\n## Description:\n\nmonitoring-aiops helps operations teams inspect and operate SolarWinds Orion, PRTG, and Zabbix monitoring environments with NOC overviews, read-only SWQL, alert rollups, health checks, and guarded maintenance actions.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, site reliability engineers, and NOC operators use this skill to triage monitoring state, answer SolarWinds SWQL questions, review PRTG and Zabbix health data, and perform audited maintenance actions across supported monitoring platforms.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Write-capable monitoring actions can acknowledge alerts, suppress monitoring, alter maintenance windows, unmanage nodes, or remove nodes without an enforced skill-side approval gate.\n\nMitigation: Use read-only monitoring accounts by default, require external change approval before write-capable credentials are enabled, and rely on dry-run, double confirmation, time-boxing, audit logs, and undo where available.\n\nRisk: Credential, TLS, and install hardening gaps could expose monitoring secrets or run an unreviewed release.\n\nMitigation: Enable TLS certificate verification in production, avoid secrets in command-line arguments or shell history, protect ~/.monitoring-aiops, and pin the package or plugin to a reviewed version.\n\nRisk: Behavior is documented as tested against mocked SolarWinds, PRTG, and Zabbix responses and still needing live NOC verification.\n\nMitigation: Run doctor and limited live checks against non-production or low-impact targets before relying on the skill for operational changes.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/monitoring-aiops)\n- [Project homepage](https://github.com/AIops-tools/Monitoring-AIops)\n- [Capabilities reference](artifact/references/capabilities.md)\n- [CLI reference](artifact/references/cli-reference.md)\n- [Setup guide](artifact/references/setup-guide.md)\n- [Agent guardrails](artifact/references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands and monitoring-operation recommendations]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs may include monitoring observations, triage steps, bounded query results, setup instructions, and proposed guarded write actions.]\n\n## Skill Version(s):\n\n0.10.2 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.10.1: 7 files, 19518 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (3229b), SKILL.md (16993b), _meta.json (136b)\n\nFile v0.10.1:SKILL.md\n\n---\nname: monitoring-aiops\nslug: monitoring-aiops\ndisplayName: \"Monitoring AIops\"\nsummary: \"Governed SolarWinds Orion + PRTG + Zabbix ops: SWQL, alert rollup, health, 42 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Monitoring-AIops\ntags: [aiops, mcp, governance, monitoring]\ndescription: >\n  Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window).\n  Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring.\n  Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill.\n  Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.\ninstaller:\n  kind: uv\n  package: monitoring-aiops\nargument-hint: \"[node/sensor id, a SWQL question, or describe your NOC task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"monitoring-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"MONITORING_AIOPS_CONFIG\",\"MONITORING_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Monitoring-AIops\",\"emoji\":\"📡\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed monitoring operations across SolarWinds Orion (SWIS REST + SWQL, port 17774 on Orion 2023.1+ with an automatic one-shot fallback to the legacy 17778, HTTP Basic auth), Paessler PRTG (web API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at /api_jsonrpc.php, API token as Bearer header on 6.4+/7.x with a legacy auth-field fallback for 6.0). Each target in the config names its own platform, so one config can span all NOCs. The governance harness (audit, policy, token/runaway budget, undo, risk-tiers) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.monitoring-aiops/ (relocatable via MONITORING_AIOPS_HOME).\n  Credentials: the Orion account password (SolarWinds), the PRTG API token, or the Zabbix API token is stored ENCRYPTED in ~/.monitoring-aiops/secrets.enc (Fernet/AES-128 + scrypt-derived key) — never plaintext on disk. Run 'monitoring-aiops init' to onboard (it asks for the platform), or 'monitoring-aiops secret set <target>' to add one. The store is unlocked by a master password from MONITORING_AIOPS_MASTER_PASSWORD (non-interactive/MCP/CI) or an interactive prompt (CLI on a TTY). A legacy plaintext env var MONITORING_<TARGET_NAME_UPPER>_SECRET is still honoured as a fallback with a deprecation warning (migrate with 'monitoring-aiops secret migrate'). The secret is used for HTTP Basic auth (SolarWinds) or as the PRTG/Zabbix API token at request time and held only in memory; secrets are never logged or echoed.\n  Read-only SWQL passthrough (swql_query) is validated to accept SELECT statements only. State-changing operations pass through the @governed_tool decorator (budget guard + audit + undo recording; each tool's risk_level is recorded as a descriptive tier, not a gate). Destructive writes (unmanage_node, remove_node, zabbix_delete_maintenance) are high-risk with dry_run + double confirmation; unmanage_node records an inverse remanage undo descriptor, zabbix_delete_maintenance captures the window's FULL definition into priorState first. Suppression/maintenance writes are TIME-BOXED (mute_alerts, schedule_maintenance, schedule_maintenance_prtg, zabbix_create_maintenance require an end time / duration). mute_alerts→unmute, pause_sensor→resume, zabbix_create_maintenance→delete-that-maintenance-id record inverse undo descriptors. Zabbix item history is BOUNDED (capped window + point count).\n  Webhooks: none — no outbound network calls beyond the configured SolarWinds SWIS / PRTG web API / Zabbix JSON-RPC endpoint.\n  SSL: verify_ssl defaults to false-friendly for self-signed lab certs; enable for production.\n  Transitive dependencies: httpx (HTTP client) and the MCP SDK. No post-install scripts or background services.\n  Validation status: behaviour is exercised against mocked SWIS/PRTG/Zabbix responses; not yet run against a live NOC (see docs/VERIFICATION.md). PRTG has a free perpetual 100-sensor Freeware edition with the API, and Zabbix is fully open source (a Docker-compose appliance is a 10-minute live check) — the easiest live checks; SolarWinds is a 30-day trial (mock-only past that — largest verification debt).\n---\n\n# Monitoring AIops\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with, endorsed by, or sponsored by SolarWinds, Paessler, Zabbix, or any monitoring vendor.** SolarWinds, Orion, SWQL, THWACK, PRTG, Paessler and Zabbix are trademarks of their respective owners. Source at [github.com/AIops-tools/Monitoring-AIops](https://github.com/AIops-tools/Monitoring-AIops) under the MIT license.\n\nGoverned network / infrastructure monitoring operations — **42 MCP tools**\nacross **SolarWinds Orion** (SWIS REST + SWQL), **Paessler PRTG** (web API),\nand **Zabbix 6.x/7.x** (JSON-RPC 2.0),\nevery one wrapped with the bundled `@governed_tool` harness: a local unified\naudit log under `~/.monitoring-aiops/`, policy engine, token/runaway budget\nguard, undo-token recording, and risk-tier labelling on the audit trail. One\nconfig can span all NOCs. The Orion password / PRTG API token / Zabbix API token is stored\n**encrypted** (`~/.monitoring-aiops/secrets.enc`, Fernet + scrypt) — never\nplaintext on disk.\n\n> **Standalone**: the governance harness is bundled in the package\n> (`monitoring_aiops.governance`) — no external skill-family dependency.\n> PRTG's free Freeware edition and an open-source Zabbix appliance are the\n> easiest live checks; SolarWinds is trial-only past 30 days (largest\n> verification debt — see `docs/VERIFICATION.md`).\n\n## What This Skill Does\n\n| Group | Platform | Tools | Count | R/W |\n|-------|----------|-------|:-----:|:---:|\n| **SWQL** | SolarWinds | library, canned, query (SELECT-only passthrough) | 3 | read |\n| **Alerts** | all | active_alerts (dedup/rollup), alert_acknowledge | 2 | 1 read, 1 write |\n| **SolarWinds health** | SolarWinds | node/nodes/interface/volume/application status, topn, noc_rollup | 7 | read |\n| **SolarWinds writes** | SolarWinds | list_events/unmanaged/muted | 3 | read |\n| | SolarWinds | mute/unmute, schedule_maintenance, remanage_node | 4 | write (med) |\n| | SolarWinds | unmanage_node, remove_node | 2 | write (**high**) |\n| **PRTG** | PRTG | sensors/sensor_details/devices/groups/history/system_status/alarms | 7 | read |\n| **PRTG writes** | PRTG | pause_sensor, resume_sensor, schedule_maintenance_prtg | 3 | write (med) |\n| **Zabbix** | Zabbix | zabbix_problems/hosts/hostgroups/triggers/events/item_history/maintenances | 7 | read |\n| **Zabbix writes** | Zabbix | zabbix_create_maintenance (time-boxed; undo = delete that id) | 1 | write (med) |\n| | Zabbix | zabbix_delete_maintenance (priorState = full definition) | 1 | write (**high**) |\n| **Undo** | all | undo_list, undo_apply | 2 | undo |\n\nThe canned SWQL library (`swql_library` lists them) answers the most-repeated\nTHWACK questions directly: `nodes_down`, `flapping_interfaces`, `muted_report`,\n`high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled`. For anything else,\n`swql_query` is a validated read-only (SELECT-only) SWQL passthrough.\n\n## Quick Install\n\n```bash\nuv tool install monitoring-aiops\nmonitoring-aiops init       # wizard: pick platform (solarwinds/prtg/zabbix) + encrypted secret\nmonitoring-aiops doctor\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@aiops-tools/monitoring-aiops\nopenclaw skills info monitoring-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- Get a NOC snapshot (`overview` / `noc_rollup`): active/unacked alert counts,\n  down/warning nodes, worst CPU\n- Answer a repeated SWQL question (`swql_library` → `swql_canned nodes_down`), or\n  run an ad-hoc read-only SWQL SELECT (`swql_query`)\n- Triage an alert storm (`active_alerts` dedup/rollup collapses flap/down\n  storms), then `alert_acknowledge`\n- SolarWinds health: `node_status`, `interface_status` (top-N by util),\n  `volume_status`, `application_status` (SAM), `topn` (cpu/mem/latency/loss)\n- PRTG: list `prtg_sensors` / `prtg_devices` / `prtg_groups`, drill with\n  `prtg_sensor_details` / `prtg_history`, check `prtg_alarms` / `prtg_system_status`\n- Zabbix: triage `zabbix_problems` (0-5 severity mapped to levels) /\n  `zabbix_triggers`, inventory `zabbix_hosts` / `zabbix_hostgroups`, drill with\n  `zabbix_item_history` (bounded), review `zabbix_events` / `zabbix_maintenances`\n- Safely take a node out for maintenance (`schedule_maintenance` /\n  `unmanage_node` with dry_run + double-confirm), pause a PRTG sensor\n  (`pause_sensor`), or create a time-boxed Zabbix maintenance window\n  (`zabbix_create_maintenance` — undo deletes exactly that window)\n\n**Do NOT use when** the target is not a SolarWinds/PRTG/Zabbix monitoring\nplatform — route hypervisor, storage, backup, cluster, network-device-config,\nor OT/industrial work to the appropriate other AIops-tools skill.\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| SolarWinds Orion / SWQL, PRTG, or Zabbix monitoring ops | **monitoring-aiops** (this skill) |\n| A non-monitoring platform (hypervisor, storage, backup, cluster, network config, OT edge) | the appropriate **other AIops-tools** skill |\n| Other monitoring stacks (not SolarWinds/PRTG/Zabbix) | out of scope for this tool |\n\n## Common Workflows\n\n> **No authorization gate**: the skill runs the operations you ask for and audits every one; it does not decide whether a write is permitted — that is the agent's judgement or the permissions of the SolarWinds/PRTG/Zabbix account it connects with (a read-only monitoring account makes writes fail at the server). There is no read-only switch, policy file, or approval gate. `MONITORING_AUDIT_APPROVED_BY` / `MONITORING_AUDIT_RATIONALE` are optional audit annotations, recorded when set.\n\n### 1. The 3 a.m. alert storm — collapse it, then acknowledge what matters\n\n1. `monitoring-aiops doctor` → confirm the NOC platform is actually reachable\n   (a \"storm\" is sometimes just a poller that lost the target)\n2. `monitoring-aiops overview` → the one-screen picture: down/warning counts\n   across the configured targets\n3. `monitoring-aiops alert list` (MCP: `active_alerts`) → deduped / rolled-up\n   entries; an interface-flap or node-down storm collapses into **one** entry\n   with a count instead of a wall of alerts\n4. `noc_rollup` → confirm whether the storm has a single upstream cause (one\n   node down taking its children with it) rather than N independent faults\n5. Acknowledge only the rolled-up entry that matters:\n   `monitoring-aiops alert ack <alert-id>` (SolarWinds `AlertActive.Acknowledge`\n   / PRTG `acknowledgealarm` / Zabbix `event.acknowledge`) — the prior ack state\n   is captured into priorState, and the ack is double-confirmed\n6. **Failure branch**: if `doctor` fails, do **not** acknowledge anything — you\n   would be silencing alerts you cannot currently see. Fix credentials with\n   `monitoring-aiops secret set <target>` first. If you acknowledged the wrong\n   alert, `monitoring-aiops undo list` → `undo apply <id>` restores the prior\n   ack state.\n\n### 2. \"Which nodes are down and what's saturated?\" (read-only)\n\n1. `noc_rollup` → down / warning counts plus the worst-CPU nodes in a single\n   call, so you do not page through a dashboard\n2. `topn cpu` (also `memory`, `latency`, `packetloss`) → the worst offenders\n   with the measured number\n3. `node_status <node>` → drill into one node; `interface_status` for a\n   suspected link problem, `volume_status` for a filling disk,\n   `application_status` for an app-layer fault\n4. `list_events` → what changed around the time things went bad\n5. `list_unmanaged` → check whether a \"missing\" node is simply unmanaged from a\n   previous maintenance window that was never reverted\n6. **Failure branch**: if a node shows down but is reachable from your shell,\n   the fault is in polling, not the node — check `list_muted` and\n   `list_unmanaged` before escalating to the network team.\n\n### 3. Planned maintenance: suppress noise time-boxed, then restore\n\n1. `node_status <node>` / `swql_canned nodes_down` → confirm you have the right\n   node and that it is currently healthy (so you can tell the difference\n   afterwards)\n2. Prefer the **time-boxed** path — it expires on its own:\n   `schedule_maintenance <node> --end ...` (SolarWinds),\n   `schedule_maintenance_prtg` (PRTG), or `zabbix_create_maintenance` (Zabbix,\n   undo → delete that maintenance id)\n3. If you genuinely need to unmanage instead:\n   `unmanage_node <node> --dry-run`, then re-run without `--dry-run` →\n   **high** risk, double confirmation; it records an inverse `remanage_node`\n   undo descriptor\n4. For a single noisy sensor rather than a whole node: `pause_sensor` (PRTG,\n   undo → `resume_sensor`) or `mute_alerts` (undo → `unmute_alerts`)\n5. When maintenance ends: `remanage_node <node>` / `resume_sensor` /\n   `unmute_alerts`, or simply `monitoring-aiops undo apply <id>` to replay the\n   recorded inverse\n6. **Failure branch**: the classic failure here is *forgetting to restore* —\n   run `list_unmanaged` and `list_muted` at the end of every maintenance window;\n   anything still listed is silently unmonitored. Time-boxed maintenance windows\n   are preferred precisely because they f\n\nArchive v0.10.0: 7 files, 19222 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2847b), SKILL.md (16673b), _meta.json (136b)\n\nArchive v0.9.0: 7 files, 19242 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2726b), SKILL.md (16790b), _meta.json (135b)\n\nArchive v0.8.0: 7 files, 19147 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2735b), SKILL.md (16790b), _meta.json (135b)\n\nArchive v0.7.0: 7 files, 19167 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2790b), SKILL.md (16790b), _meta.json (135b)\n\nArchive v0.6.0: 7 files, 19058 bytes\n\nFiles: references/agent-guardrails.md (7220b), references/capabilities.md (6422b), references/cli-reference.md (2803b), references/setup-guide.md (5337b), skill-card.md (2457b), SKILL.md (16790b), _meta.json (135b)\n\nArchive v0.5.0: 7 files, 18916 bytes\n\nFiles: references/agent-guardrails.md (6856b), references/capabilities.md (6422b), references/cli-reference.md (2774b), references/setup-guide.md (5380b), skill-card.md (2709b), SKILL.md (16540b), _meta.json (135b)","readmeExcerpt":"Skill: monitoring-aiops Owner: zw008 Summary: Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install monitoring-aiops\nmonitoring-aiops init       # wizard: pick platform (solarwinds/prtg/zabbix) + encrypted secret\nmonitoring-aiops doctor"},{"language":"bash","snippet":"openclaw plugins install clawhub:@zw008/monitoring-aiops\nopenclaw skills info monitoring-aiops          # expect: Visible to model: yes"},{"language":"text","snippet":"You operate an enterprise monitoring system through the monitoring-aiops MCP\ntools. A target is SolarWinds Orion, Paessler PRTG, or Zabbix — check which\nbefore reasoning about what a field means.\n\nTOOL USE\n- Before answering any question about the current monitored estate, you MUST\n  call a tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer. A read that fails returns an \"error\" field rather\n  than raising — treat that as \"unknown\", not as \"everything is fine\".\n- Prefer the named canned SWQL (swql_library / swql_canned) over writing your\n  own query. Hand-written SWQL is where small models most often produce\n  syntactically valid nonsense.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete. An alert count from a truncated read is not\n  the number of alerts.\n- A null field means the platform did not return that value. Report it as \"not\n  available\" — never infer it.\n- Report values exactly as returned. Do not translate status codes, severity\n  names, node captions, or sensor names into your own vocabulary.\n- Acknowledged is not the same as resolved. An acknowledged alert is still\n  active; say which you mean.\n\n- None of the node or maintenance-window writes has a CLI command, so nothing will ask you\n  to confirm them — `remove_node` and `unmanage_node` least of all. Call with\n  `dry_run=True` first and wait for an explicit go-ahead.\n\nSCOPE\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert an outage, root cause, or business impact unless a tool result\n  supports it. A down sensor is one s"},{"language":"bash","snippet":"monitoring-aiops doctor"},{"language":"bash","snippet":"export MONITORING_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport MONITORING_AUDIT_RATIONALE=\"change window CHG0041231, muting core-sw1\""},{"language":"bash","snippet":"monitoring-aiops init                      # interactive wizard (asks for the platform: solarwinds/prtg/zabbix)\nmonitoring-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   SolarWinds: a SWQL query · PRTG: /api/status.json\n                                           #   Zabbix: apiinfo.version (no auth) + authed host count\nmonitoring-aiops mcp                       # start the MCP server (stdio transport)"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: monitoring-aiops\nslug: monitoring-aiops\ndisplayName: \"Monitoring AIops\"\nsummary: \"Governed SolarWinds Orion + PRTG + Zabbix ops: SWQL, alert rollup, health, 42 tools.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/Monitoring-AIops\ntags: [aiops, mcp, governance, monitoring]\ndescription: >\n  Use this skill whenever the user needs to operate a network / infrastructure monitoring NOC on SolarWinds Orion (SWIS REST + SWQL), Paessler PRTG (web API), or Zabbix 6.x/7.x (JSON-RPC) — a one-shot NOC overview, canned SWQL answers (nodes down, flapping interfaces, muted, high-CPU nodes, full volumes, unmanaged/scheduled), a validated read-only SWQL passthrough, deduped/rolled-up active alerts, SolarWinds node/interface/volume/application health and top-N, PRTG sensors/devices/groups/history/alarms, Zabbix problems/hosts/host-groups/triggers/events/item-history/maintenances, and guarded writes (acknowledge, mute/unmute, schedule maintenance, unmanage/remanage, remove node, pause/resume sensor, create/delete Zabbix maintenance window).\n  Always use this skill for \"SolarWinds\", \"Orion\", \"SWQL\", \"THWACK question\", \"PRTG\", \"Paessler\", \"Zabbix\", \"Zabbix problem\", \"Zabbix trigger\", \"Zabbix maintenance\", \"NOC overview\", \"which nodes are down\", \"flapping interfaces\", \"interface flap storm\", \"alert storm\", \"acknowledge this alert\", \"worst CPU nodes\", \"top-N by latency/packet loss\", \"which volumes are full\", \"muted alerts report\", \"unmanaged nodes\", \"schedule a maintenance window\", \"unmanage / remanage a node\", \"pause a PRTG sensor\" when the context is monitoring.\n  Do NOT use when the target is something other than a SolarWinds/PRTG/Zabbix monitoring platform (a hypervisor, storage appliance, backup product, Kubernetes cluster, network device config, or OT/industrial equipment) — route those to the appropriate other AIops-tools skill.\n  Governed monitoring operations with a built-in governance harness (audit, policy, token budget, undo, risk-tiers). PRTG's free Freeware edition and an open-source Zabbix appliance are the easiest live checks; SolarWinds is trial-only past 30 days.\ninstaller:\n  kind: uv\n  package: monitoring-aiops\nargument-hint: \"[node/sensor id, a SWQL question, or describe your NOC task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"monitoring-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"MONITORING_AIOPS_CONFIG\",\"MONITORING_AIOPS_MASTER_PASSWORD\"]},\"homepage\":\"https://github.com/AIops-tools/Monitoring-AIops\",\"emoji\":\"📡\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed monitoring operations across SolarWinds Orion (SWIS REST + SWQL, port 17774 on Orion 2023.1+ with an automatic one-shot fallback to the legacy 17778, HTTP Basic auth), Paessler PRTG (web API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at /api_jsonrpc.php, API token as Bearer header on 6.4+/7.x with a legacy auth-field fallback for 6.0). Each target in the config names its own platform, so one config can span all NOCs"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"monitoring-aiops\",\n  \"version\": \"0.10.4\",\n  \"publishedAt\": 1789601158731\n}"},{"path":"references/agent-guardrails.md","content":"# Agent guardrails — running monitoring-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The account you connect with.** Give it a SolarWinds/PRTG/Zabbix login with\n  read-only monitoring scope. A write then fails at the server, which is the\n  only place the permission actually lives — no skill-side flag can be argued\n  around by a model, but a revoked permission cannot be.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Never write to the monitoring database\" | `swql_query` accepts a single read-only `SELECT` and nothing else — no verb invoke, no multi-statement, no DELETE. Orion state changes only happen through the named, governed write tools. |\n| \"Don't invent a value when a field is missing\" | A column the platform did not return comes back as `null`, never as `\"\"`. An absent Orion `StatusDescription`, a PRTG sensor `message`, a Zabbix host `dns` — all distinguishable from a genuinely empty one. |\n| \"Tell me if the output was cut off\" | Every row-capped read returns `returned` / `limit` / `truncated`: `swql_query`, `swql_canned`, `list_events`, `zabbix_events`, `zabbix_item_history`, and `interface_status` with a `top`. Truncation is measured — one row past the cap is fetched, or the full set is counted before the cut — never guessed from a length coincidence. |\n| \"Deduplicate the alert storm before showing me\" | `active_alerts` already rolls repeats of the same message into one row with a `count` and up to three `examples`, worst-first. Report the rollup; do not re-count the raw list. |\n| \"Normalise severity across platforms\" | Zabbix's 0–5 scale is already mapped to canonical `level` values (`info`/`warning`/`high`/`critical`) alongside the platform's own `severity` name. Use `level` for cross-platform statements and `severity` when quoting the platform. |\n| \"Confirm before anything disruptive\" | `remove_node`, `unmanage_node`, `mute_alerts` and the main"},{"path":"references/capabilities.md","content":"# monitoring-aiops capabilities\n\n> **42 MCP tools** (30 read, 10 write, 2 undo) across SolarWinds\n> Orion (SWIS REST + SWQL, port 17774 with a legacy-17778 fallback, HTTP\n> Basic auth), Paessler PRTG (web\n> API, port 443/8080, API token), and Zabbix 6.x/7.x (JSON-RPC 2.0 at\n> `/api_jsonrpc.php`, API token — Bearer header on 6.4+/7.x, legacy `auth`\n> field fallback for 6.0). Each config target names its own `platform`.\n> SWIS/PRTG/Zabbix responses are mocked and need live verification.\n\n## SWQL — SolarWinds (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `swql_library` | (local) | the catalogue of canned queries: `nodes_down`, `flapping_interfaces`, `muted_report`, `high_cpu_nodes`, `volumes_full`, `unmanaged_scheduled` |\n| `swql_canned` | named SWQL → SWIS `/Query` | rows for the named canned query |\n| `swql_query` | validated read-only SWQL → SWIS `/Query` | rows for a caller SELECT (SELECT-only; rejected otherwise) |\n\n## Alerts — all platforms\n\n| Tool | Risk | Path | Returns / effect |\n|------|------|------|------------------|\n| `active_alerts` | read | SWIS `AlertActive`/`AlertObjects`, PRTG `/api/table.json?content=messages`, or Zabbix `problem.get` | active alerts **deduped/rolled up by message** — flap/down storms collapse into one counted entry |\n| `alert_acknowledge` | write **medium** | SW `AlertActive.Acknowledge` verb / PRTG `acknowledgealarm.htm` / Zabbix `event.acknowledge` (action 6; prior ack state → priorState) | acknowledges an alert / alarm / problem event |\n\n## SolarWinds health (read)\n\n| Tool | SWQL / path | Returns |\n|------|-------------|---------|\n| `node_status` | `Orion.Nodes` | one node's status, CPU/mem, response time |\n| `nodes_list` | `Orion.Nodes` | node inventory (status, vendor, IP, last boot) |\n| `interface_status` | `Orion.NPM.Interfaces` | top-N interfaces by utilisation (in/out, errors, oper status) |\n| `volume_status` | `Orion.Volumes` | volumes by % used (size, used, type) |\n| `application_status` | `Orion.APM.Application` (SAM) | SAM application/component status |\n| `topn` | `Orion.Nodes` metrics | top-N nodes by `cpu` / `memory` / `latency` / `packetloss` |\n| `noc_rollup` | folds `Orion.Nodes` | down/warning counts + worst-CPU nodes in one call |\n\n## SolarWinds writes\n\n| Tool | Risk | Path / verb | Undo / safety |\n|------|------|-------------|---------------|\n| `list_events` | read | `Orion.Events` | recent events (read) |\n| `list_unmanaged` | read | `Orion.Nodes` where Unmanaged | currently-unmanaged nodes (read) |\n| `list_muted` | read | `Orion.AlertSuppression` | currently-muted objects (read) |\n| `mute_alerts` | write **med** | `AlertSuppression` (SuppressAlerts) | **time-boxed** (requires end time); records inverse **unmute** undo |\n| `unmute_alerts` | write **med** | `AlertSuppression` (ResumeAlerts) | un-suppresses alerting |\n| `schedule_maintenance` | write **med** | `AlertSuppression` window | **requires an end time** (time-boxed maintenance window) |\n| `unmanage_node` |"},{"path":"references/cli-reference.md","content":"# monitoring-aiops CLI reference\n\n> Covers SolarWinds Orion (SWIS REST + SWQL), Paessler\n> PRTG (web API), and Zabbix 6.x/7.x (JSON-RPC); SWIS/PRTG/Zabbix responses are\n> mocked and need live verification.\n> The CLI is a convenience subset — the full 42-tool surface is via the MCP\n> server (`monitoring-aiops mcp`).\n\n## Setup & diagnostics\n\n```bash\nmonitoring-aiops init                      # interactive wizard (asks for the platform: solarwinds/prtg/zabbix)\nmonitoring-aiops doctor [--skip-auth]      # config + secret store + connectivity\n                                           #   SolarWinds: a SWQL query · PRTG: /api/status.json\n                                           #   Zabbix: apiinfo.version (no auth) + authed host count\nmonitoring-aiops mcp                       # start the MCP server (stdio transport)\n```\n\n## Secrets (encrypted store ~/.monitoring-aiops/secrets.enc)\n\n```bash\nmonitoring-aiops secret set <target> [--value <secret>]  # store Orion password / PRTG or Zabbix token (hidden prompt if no --value)\nmonitoring-aiops secret list                             # names only — secrets never shown\nmonitoring-aiops secret rm <target>\nmonitoring-aiops secret migrate                          # import legacy plaintext env (MONITORING_<TARGET>_SECRET)\nmonitoring-aiops secret rotate-password                  # re-encrypt under a new master password\n```\n\n## Overview\n\n```bash\nmonitoring-aiops overview [--target <t>]   # NOC summary: platform + active/unacked alert counts + top rollup\n```\n\n## SWQL (SolarWinds)\n\n```bash\nmonitoring-aiops swql library                    # list the canned queries\nmonitoring-aiops swql canned <name>              # run a canned query: nodes_down, flapping_interfaces,\n                                                 #   muted_report, high_cpu_nodes, volumes_full, unmanaged_scheduled\nmonitoring-aiops swql query \"SELECT ...\"         # validated read-only SWQL passthrough (SELECT only)\n```\n\n## Alerts (all platforms)\n\n```bash\nmonitoring-aiops alert list [--target <t>]       # active alerts, deduped/rolled up by message\nmonitoring-aiops alert ack <alert_id>            # acknowledge an alert / PRTG alarm / Zabbix problem event\n```\n\n## Common options\n\n- `--target, -t <name>` — target name from `config.yaml` (omit to use the\n  default/first target); each target declares its own `platform`\n- `overview`, `swql`, and `alert` are the CLI subset; the remaining SolarWinds\n  health, PRTG, Zabbix, and governed-write tools (mute/unmute,\n  schedule_maintenance, unmanage/remanage/remove node, PRTG pause/resume,\n  Zabbix maintenance create/delete) are exposed through the MCP server.\n  High-risk MCP writes use dry-run + double-confirm; `MONITORING_AUDIT_APPROVED_BY`\n  / `MONITORING_AUDIT_RATIONALE` are recorded on the audit row when set."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2202,"uniquenessScore":38,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T04:22:24.953Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T09:09:08.206Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}