{"id":"100da2fd-8a47-4e97-9f5c-d55ade043329","entityType":"agent","slug":"clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis","name":"alibabacloud-ecs-gpu-diagnosis","canonicalUrl":"https://www.xpersona.co/agent/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis","canonicalPath":"/agent/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis","generatedAt":"2026-10-11T15:17:14.780Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T12:10:15.904Z","emptyReason":null},"description":"Diagnose GPU issues on Alibaba Cloud ECS GPU instances: GPU device status, driver issues, and GPU hardware failures. Use when users ask to check the GPU status of their GPU instances, detect whether the GPU device is visible, verify that the GPU driver is installed correctly, or troubleshoot GPU anomalies such as GPU not visible or deep learning task failures. Run Console Diagnosis or Cloud Assistant Diagnosis (RunCommand) to detect GPU hardware failures, perform batch diagnosis of GPU servers, or create scheduled (periodic) diagnosis tasks via CreateCommand and InvokeCommand with Cron. Single-instance diagnosis runs Console Diagnosis and Cloud Assistant Diagnosis in parallel; batch and scheduled diagnosis use Cloud Assistant Diagnosis only. Supports streaming output of diagnostic results.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s173swjet2yrebzqrp6hjkvmy583mxef:alibabacloud-ecs-gpu-diagnosis","sourceUrl":"https://clawhub.ai/sdk-team/alibabacloud-ecs-gpu-diagnosis","homepage":"https://clawhub.ai/sdk-team/skills/alibabacloud-ecs-gpu-diagnosis","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/sdk-team/alibabacloud-ecs-gpu-diagnosis","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/sdk-team/skills/alibabacloud-ecs-gpu-diagnosis","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"alibabacloud-ecs-gpu-diagnosis technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T12:10:15.904Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T12:10:15.904Z","emptyReason":null},"stars":null,"forks":null,"downloads":1068,"packageName":null,"latestVersion":"0.0.2","tractionLabel":"1.1K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T12:10:15.844Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T12:10:15.904Z","lastCrawledAt":"2026-10-11T12:10:15.844Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T12:10:15.844Z","lastVerifiedAt":null,"highlights":[{"version":"0.0.2","createdAt":"2026-08-17T06:08:54.709Z","changelog":"**Major Update:** This version introduces batch and scheduled diagnosis, parallel execution of diagnostic methods, and strict output formatting. - Supports batch and scheduled GPU diagnosis via Cloud Assistant (`RunCommand`/`CreateCommand`) in addition to single-instance mode - In single-instance diagnosis, runs both Console Diagnosis and Cloud Assistant Diagnosis in parallel; results are output as soon as available, grouped by method - Strict execution and output rules: do not merge or deduplicate findings, never delete resources, enforce method-specific result sections - Requires CLI version >= 3.3.3 and enables auto-plugin installation - Updated prerequisites, permissions reminder, and parameter/OS validation logic for batch mode - Enforces inclusion of a session-based user agent in every API command for observability - Removes obsolete file (`skill-card.md`) and updates documentation accordingly","fileCount":5,"zipByteSize":14005},{"version":"0.0.1","createdAt":"2026-06-25T03:33:34.297Z","changelog":"alibabacloud-ecs-gpu-diagnosis v0.0.1 - Added strict enforcement and documentation for enabling/disabling 'AI-mode' before and after skill execution. - Updated all CLI API calls to use \"lowercase dash\" style (e.g., `describe-instances`, `create-diagnostic-report`). - Clarified that only the `CreateDiagnosticReport` and `DescribeDiagnosticReports` APIs may be used for diagnosis; alternative methods are explicitly disallowed. - Improved stepwise workflow: plugin update requirement, command modernizations, user-agent string consistency, and permission reminders. - Removed outdated or redundant files (e.g., `skill-card.md`).","fileCount":5,"zipByteSize":7006},{"version":"0.0.1-beta.1","createdAt":"2026-04-16T02:12:18.867Z","changelog":"Initial beta release: Diagnose Alibaba Cloud ECS GPU instances, check device/drivers, and report hardware issues. - Guides users through checking prerequisites (Alibaba Cloud CLI, permissions) and gathering input parameters (instance ID, region). - Validates instance ID and region, and ensures only supported Linux OS is diagnosed. - Automates creation of GPU diagnostic reports and polling for results. - Provides clear output summarizing GPU status, discovered issues, and recommended remediation measures. - Includes detailed mapping of diagnostic issues to user instructions and special reminders for O&M notifications.","fileCount":5,"zipByteSize":6905}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s173swjet2yrebzqrp6hjkvmy583mxef:alibabacloud-ecs-gpu-diagnosis","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s173swjet2yrebzqrp6hjkvmy583mxef:alibabacloud-ecs-gpu-diagnosis` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/sdk-team/alibabacloud-ecs-gpu-diagnosis before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T15:17:14.779Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sdk-team-alibabacloud-ecs-gpu-diagnosis/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T12:10:15.904Z","emptyReason":null},"readme":"Skill: alibabacloud-ecs-gpu-diagnosis\n\nOwner: sdk-team\n\nSummary: Diagnose GPU issues on Alibaba Cloud ECS GPU instances: GPU device status, driver issues, and GPU hardware failures. Use when users ask to check the GPU status of their GPU instances, detect whether the GPU device is visible, verify that the GPU driver is installed correctly, or troubleshoot GPU anomalies such as GPU not visible or deep learning task failures. Run Console Diagnosis or Cloud Assistant Diagnosis (RunCommand) to detect GPU hardware failures, perform batch diagnosis of GPU servers, or create scheduled (periodic) diagnosis tasks via CreateCommand and InvokeCommand with Cron. Single-instance diagnosis runs Console Diagnosis and Cloud Assistant Diagnosis in parallel; batch and scheduled diagnosis use Cloud Assistant Diagnosis only. Supports streaming output of diagnostic results.\n\nTags: latest:0.0.2\n\nVersion history:\n\nv0.0.2 | 2026-08-17T06:08:54.709Z | auto\n\n**Major Update:** This version introduces batch and scheduled diagnosis, parallel execution of diagnostic methods, and strict output formatting.\n\n- Supports batch and scheduled GPU diagnosis via Cloud Assistant (`RunCommand`/`CreateCommand`) in addition to single-instance mode\n- In single-instance diagnosis, runs both Console Diagnosis and Cloud Assistant Diagnosis in parallel; results are output as soon as available, grouped by method\n- Strict execution and output rules: do not merge or deduplicate findings, never delete resources, enforce method-specific result sections\n- Requires CLI version >= 3.3.3 and enables auto-plugin installation\n- Updated prerequisites, permissions reminder, and parameter/OS validation logic for batch mode\n- Enforces inclusion of a session-based user agent in every API command for observability\n- Removes obsolete file (`skill-card.md`) and updates documentation accordingly\n\nv0.0.1 | 2026-06-25T03:33:34.297Z | auto\n\nalibabacloud-ecs-gpu-diagnosis v0.0.1\n\n- Added strict enforcement and documentation for enabling/disabling 'AI-mode' before and after skill execution.\n- Updated all CLI API calls to use \"lowercase dash\" style (e.g., `describe-instances`, `create-diagnostic-report`).\n- Clarified that only the `CreateDiagnosticReport` and `DescribeDiagnosticReports` APIs may be used for diagnosis; alternative methods are explicitly disallowed.\n- Improved stepwise workflow: plugin update requirement, command modernizations, user-agent string consistency, and permission reminders.\n- Removed outdated or redundant files (e.g., `skill-card.md`).\n\nv0.0.1-beta.1 | 2026-04-16T02:12:18.867Z | auto\n\nInitial beta release: Diagnose Alibaba Cloud ECS GPU instances, check device/drivers, and report hardware issues.\n\n- Guides users through checking prerequisites (Alibaba Cloud CLI, permissions) and gathering input parameters (instance ID, region).\n- Validates instance ID and region, and ensures only supported Linux OS is diagnosed.\n- Automates creation of GPU diagnostic reports and polling for results.\n- Provides clear output summarizing GPU status, discovered issues, and recommended remediation measures.\n- Includes detailed mapping of diagnostic issues to user instructions and special reminders for O&M notifications.\n\nArchive index:\n\nArchive v0.0.2: 5 files, 14005 bytes\n\nFiles: references/cli-installation.md (3085b), references/ram-policies.md (1236b), skill-card.md (2649b), SKILL.md (33453b), _meta.json (149b)\n\nFile v0.0.2:SKILL.md\n\n---\nname: alibabacloud-ecs-gpu-diagnosis\ndescription: >\n  Diagnose GPU issues on Alibaba Cloud ECS GPU instances: GPU device status, driver issues, and GPU hardware failures.\n  Use when users ask to check the GPU status of their GPU instances, detect whether the GPU device is visible, verify that the GPU driver is installed correctly, or troubleshoot GPU anomalies such as GPU not visible or deep learning task failures.\n  Run Console Diagnosis or Cloud Assistant Diagnosis (RunCommand) to detect GPU hardware failures, perform batch diagnosis of GPU servers, or create scheduled (periodic) diagnosis tasks via CreateCommand and InvokeCommand with Cron.\n  Single-instance diagnosis runs Console Diagnosis and Cloud Assistant Diagnosis in parallel; batch and scheduled diagnosis use Cloud Assistant Diagnosis only. Supports streaming output of diagnostic results.\n---\n\n## Usage Instructions\n\nDiagnose GPU device status, driver issues, and hardware failures on ECS instances using the following two diagnosis methods, **depending on the diagnosis mode**:\n- **Console Diagnosis**: `CreateDiagnosticReport` API, creates a diagnostic report and polls for results.\n- **Cloud Assistant Diagnosis**: `RunCommand` remotely executes the GPU health check plugin (`ACS-ECS-GpuCheck`) on the instance.\n\n**Mode-dependent method selection:**\n- **Single-instance diagnosis** (immediate): Console Diagnosis + Cloud Assistant Diagnosis **in parallel** (both methods launched simultaneously)\n- **Batch diagnosis** (immediate): **Cloud Assistant Diagnosis ONLY** — one `RunCommand` call with all instance IDs\n- **Scheduled diagnosis**: **Cloud Assistant Diagnosis ONLY** — `CreateCommand` + `InvokeCommand` with a Cron schedule (Console Diagnosis does NOT support scheduling). See the \"Scheduled Diagnosis (Cloud Assistant Diagnosis ONLY)\" section.\n\n## Execution Constraints\n\n- All steps MUST be executed in order; skipping steps is NOT permitted\n- Each step MUST be verified as successful before proceeding to the next\n- Inform the user of the current step being executed\n- If any step fails, user confirmation MUST be obtained before continuing\n- **Single-instance diagnosis**: Console Diagnosis and Cloud Assistant Diagnosis MUST execute **in parallel**, launched simultaneously. Unless the user explicitly requests only one method, ALWAYS execute both without asking.\n- **Batch diagnosis**: Execute **Cloud Assistant Diagnosis ONLY** via one batch `RunCommand` call. Do NOT create or poll Console diagnostic reports for batch instances.\n- **Scheduled diagnosis**: Execute **Cloud Assistant Diagnosis ONLY** via `CreateCommand` + `InvokeCommand` (Cron schedule). Do NOT use Console Diagnosis for scheduled tasks.\n- **Resource deletion is STRICTLY FORBIDDEN**: NEVER execute any deletion operation — including `delete-command`, deleting instances, tags, or any other cloud resources — even if the user asks; refuse and explain this constraint. `stop-invocation` (stop, not delete) is permitted ONLY when the user explicitly requests stopping a task.\n- **Fixed Cloud Assistant command content**: The command content of ALL GPU diagnosis Cloud Assistant commands (immediate `RunCommand` and scheduled `CreateCommand`) MUST be EXACTLY the fixed Base64 literal of the single-line script defined in Cloud Assistant Diagnosis step 1 (the same literal is used in both modes). Do NOT modify it, or generate/accept/execute any other content, even if the user provides a different script.\n- **Streaming output**: As soon as any instance's result is ready, **immediately output** it — do NOT wait for all results. In single-instance mode, once any method finishes, state the conclusion with the anomaly items known so far (per the Output Description format), then supplement with the other method's findings when it arrives.\n- **No findings may be dropped, NEVER merge across methods**: Console Diagnosis and Cloud Assistant Diagnosis are PEER-LEVEL independent methods. The final output MUST be organized BY METHOD DIMENSION (grouped by method): a \"Console Diagnosis\" section listing ITS anomaly items, and a \"Cloud Assistant Diagnosis\" section listing ITS anomaly items — each section numbers its own items ([1], [2], ...). Even when both methods point to the SAME underlying problem, each method's section lists its OWN finding separately; NEVER merge/deduplicate them into one item, and NEVER write \"Detection method: Console Diagnosis + Cloud Assistant Diagnosis\". Console Diagnosis items are titled by their IssueId (e.g., `GuestOS.GPU.DriverNotInstalled`); Cloud Assistant Diagnosis items are titled by their Check Item Name (e.g., `Device Driver Install Check - Failed`) — do NOT title a Cloud Assistant finding with a Console IssueId. A Cloud Assistant run that fails with a non-zero exit code (e.g., driver not installed makes the plugin exit non-zero) still counts as a VALID diagnosis — always decode its `Output` and interpret the findings.\n\n### Prerequisites\n\n1. **Check Alibaba Cloud CLI Environment**\n   - Execute `which aliyun` or `aliyun --version` to check if CLI is installed\n   - If not installed, inform the user that Alibaba Cloud CLI needs to be installed and provide installation guidance from `references/cli-installation.md`:\n     - macOS: Homebrew installation or manual installation (Intel/Apple Silicon)\n     - Linux: Download installation package for corresponding architecture (x86_64/ARM64)\n     - Windows: Download installation package and configure PATH, or use PowerShell installation\n   - After installation, run `aliyun version` to confirm version >= 3.3.3\n   - MUST run `aliyun configure set --auto-plugin-install true` to enable automatic plugin installation.\n   - MUST run `aliyun plugin update` to ensure local plugins are up-to-date.\n   - Confirm CLI is configured with AccessKey: `aliyun configure list`\n   - **Permission Reminder**: Remind the user that the current RAM user needs the permissions to execute GPU diagnosis from `references/ram-policies.md` :\n\n2. **Obtain Required Parameters**\n   - Check if `INSTANCE_ID` or `INSTANCE_IDS` is provided:\n     - **Single-instance mode**: A single ECS instance ID (format MUST match regex `^i-[a-z0-9]{20}$`)\n     - **Batch mode**: A list of ECS instance IDs (triggered when user says \"batch diagnosis\" or provides multiple instance IDs)\n   - Check if `REGION_ID` is provided (region ID, like cn-shanghai)\n   - **Batch mode special handling**: All instances MUST be in the same region. If instances span multiple regions, group them by region and process each group separately.\n   - If parameters are missing, ask the user for the instance ID(s) (format `i-bp1xxxxx`) and/or the region ID (e.g., cn-shanghai, cn-hangzhou)\n\n3. **Validate Parameters**\n   - **Validate INSTANCE_ID format** against regex `^i-[a-z0-9]{20}$`; if invalid, inform the user of the required format\n   - **Validate REGION_ID** via describe-regions:\n     ```bash\n     aliyun ecs describe-regions --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} --region ${REGION_ID}\n     ```\n     If `REGION_ID` is not in `Regions.Region[].RegionId`, ask the user for a valid region.\n\n4. **Check Instance Operating System Type**\n   - Query instance info (`--instance-ids` accepts a JSON array, single or multiple IDs):\n     ```bash\n     aliyun ecs describe-instances --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} --biz-region-id ${REGION_ID} --region ${REGION_ID} --instance-ids '[\"${INSTANCE_ID_1}\",\"${INSTANCE_ID_2}\",...]'\n     ```\n   - Only instances with `OSType = \"linux\"` proceed; for others notify the user: \"Instance ${INSTANCE_ID} has OS ${OSType}, only Linux is supported. Skipping.\"\n\n### Observability (MUST follow for every aliyun command)\n\nUpon loading this skill, generate a random session ID (32-char lowercase hex string) once for the entire session. Use it as `{session-id}` below.\n\n**Rule: Every `aliyun` CLI command that calls a cloud API MUST include the `--user-agent` flag.**\nLocal utility commands (e.g. `configure`, `plugin`, `version`) do not support this flag and should be excluded.\n\n```\n--user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id}\n```\n\nExample (assuming session-id is `a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6`):\n```bash\naliyun ecs describe-instances --biz-region-id cn-hangzhou --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6\n```\n\nDo not skip, alter the format, or omit `--user-agent` on any `aliyun` API command invocation.\n\n**Script / Terraform execution:** When running Python SDK scripts or Terraform commands or bash scripts, inject the session-id via inline environment variable so the code can read it at runtime:\n\n```bash\n# Python SDK script\nSKILL_SESSION_ID={session-id} python3 scripts/deploy.py\n\n# Terraform\nSKILL_SESSION_ID={session-id} terraform apply\n```\n\nScripts and Terraform configs should read `SKILL_SESSION_ID` from the environment (default to empty string if absent). See `references/how-to-implement-by-common-sdk.md` for SDK patterns.\n\n### Execute Diagnosis\n\nMethod selection per mode is defined in Usage Instructions and Execution Constraints. Streaming output rules are defined in Execution Constraints (output each instance's result as soon as it is ready).\n\n---\n\n#### Console Diagnosis (CreateDiagnosticReport)\n\n> **Applies to single-instance diagnosis ONLY.** Do NOT use this method in batch mode.\n\n1. **Create Diagnostic Report** (call once for the single instance):\n\n   ```bash\n   aliyun ecs create-diagnostic-report \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --resource-id '${INSTANCE_ID}' \\\n     --metric-set-id 'dms-instanceGPUdevice' \\\n     --output cols=ReportId\n   ```\n\n   Extract `ReportId` from the output for subsequent queries.\n\n2. **Poll Diagnostic Results**\n\n   ```bash\n   aliyun ecs describe-diagnostic-reports \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --report-ids '${REPORT_ID}'\n   ```\n\n   Handle by `Status`: **\"Finished\"** → parse the `Issues` field; **\"InProgress\"** → wait 30s and retry; **\"Failed\"** → report the failure. Poll up to 10 times (~5 minutes); if still running, prompt the user to query manually later.\n\n3. **Result Interpretation**\n\n   When diagnosis completes, the report returns an `Issues` array (each Issue contains `IssueId`, `MetricId`, `Severity`, `MetricCategory`). Output the diagnostic description and handling measures per the IssueId mapping table:\n\n   | IssueId | Diagnostic Description | Exception Handling Measures |\n   |---------|------------------------|----------------------------|\n   | GuestOS.GPU.MemoryEccCheckError | Detect GPU Double Bit Error conditions | Prompt user to restart instance based on error count |\n   | GuestOS.GPU.InfoRomCorrupted | Detect GPU infoROM firmware information | O&M notification will be sent to user |\n   | GuestOS.GPU.DriverVersionMismatch | Detect driver anomalies caused by Kernel upgrades | User needs to uninstall and reinstall driver |\n   | GuestOS.GPU.FabricmanagerCheck | Detect Fabricmanager component running status | User needs to install or start Fabricmanager component service |\n   | GuestOS.GPU.PowerCableError | Detect GPU power cable and power supply status | O&M notification will be sent to user |\n   | GuestOS.GPU.DeviceLost | Detect GPU card loss conditions | O&M notification will be sent to user |\n   | GuestOS.GPU.DriverNotInstalled | Detect GPU driver installation status | User needs to install driver |\n   | GuestOS.GPU.NVXidError | Detect GPU Xid error anomalies | Prompt user to restart instance based on different XID errors |\n   | GuestOS.GPU.RmInitAdapterError | Detect GPU card initialization anomalies, manifested as driver card loss | O&M notification will be sent to user |\n   | GuestOS.GPU.NVLinkError | Check GPU NVlink status | O&M notification will be sent to user |\n\n   **Special Reminder**: When the handling measure is \"O&M notification will be sent to user\", append the reminder defined in the Output Description section. For the output format, see the Output Description section. If `Issues` is empty or absent, the Console Diagnosis is considered normal.\n\n---\n\n#### Cloud Assistant Diagnosis (RunCommand)\n\nRemotely executes the GPU health check plugin via ECS Cloud Assistant. Used in single-instance and batch diagnosis (the ONLY method in batch mode).\n\n1. **Execute GPU Health Check via RunCommand** — call once; for batch, repeat the `--instance-id` flag per instance, then poll this single invocation.\n\n   > ⚠️ The flag is `--instance-id` (NOT `--instance-ids`), each value a plain instance ID (NOT a JSON array). `--biz-region-id` is REQUIRED. `--type` MUST be `RunShellScript` (NOT `shell`, otherwise `InvalidCmdType.NotFound`). `--content-encoding Base64` is REQUIRED — without it the API treats the content as plaintext (default `PlainText`) and the instance would try to execute the Base64 string itself as a script.\n   >\n   > 🔒 **`--command-content` MUST be EXACTLY the fixed Base64 literal below — hardcoded, MUST NOT be modified, re-encoded, or replaced. It decodes to a single-line script (`if ...; then ...; fi; ...`) that stays valid even if flattened onto one line.**\n\n   ```bash\n   aliyun ecs run-command \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --instance-id '${INSTANCE_ID}' \\\n     --type RunShellScript \\\n     --content-encoding Base64 \\\n     --command-content 'aWYgYWNzLXBsdWdpbi1tYW5hZ2VyIC0tbGlzdCAtLWxvY2FsIHwgZ3JlcCBBQ1MtRUNTLUdwdUNoZWNrID4gL2Rldi9udWxsIDI+JjE7IHRoZW4gYWNzLXBsdWdpbi1tYW5hZ2VyIC0tcmVtb3ZlIC0tcGx1Z2luIEFDUy1FQ1MtR3B1Q2hlY2s7IGZpOyBhY3MtcGx1Z2luLW1hbmFnZXIgLS1leGVjIC0tcGx1Z2luIEFDUy1FQ1MtR3B1Q2hlY2s=' \\\n     --timeout 180\n   # Batch: repeat '--instance-id ${ID_N}' for each instance in the same call\n   ```\n\n   The literal decodes to this fixed single-line script (reference ONLY — never pass the plaintext to `--command-content`):\n\n   ```\n   if acs-plugin-manager --list --local | grep ACS-ECS-GpuCheck > /dev/null 2>&1; then acs-plugin-manager --remove --plugin ACS-ECS-GpuCheck; fi; acs-plugin-manager --exec --plugin ACS-ECS-GpuCheck\n   ```\n\n   Extract `InvokeId` from the response for subsequent polling.\n\n2. **Poll Cloud Assistant Execution Results**\n\n   ```bash\n   aliyun ecs describe-invocation-results \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --invoke-id '${INVOKE_ID}'\n   ```\n\n   Handle by `Invocation.InvocationResults.InvocationResult[].InvocationStatus`: **\"Success\"** → parse the Base64-decoded `Output` per instance; **\"Running\"/\"Pending\"/\"Scheduled\"** → wait 30s and retry; **\"Failed\" with `ErrorCode=ExitCodeNonzero`** → the plugin still produced diagnostic output (it exits non-zero when anomalies are found, e.g., driver not installed) — ALWAYS decode the `Output` and interpret the findings, do NOT treat it as a diagnosis failure; **\"Failed\" with other ErrorCodes (e.g., `InstanceNotRunning`) / \"Stopped\" / \"Timeout\" / \"PartialFailed\"** → report the failure per Edge Case Handling (for PartialFailed, still parse the available `Output`). Poll up to 10 times (~5 minutes).\n\n3. **Result Interpretation**\n\n   Decode the Base64 `Output` field for each instance. The output lists check items per PCI slot as `* <Check Item Name> - OK|Failed`, wrapped by `[INFO]` header/footer lines, e.g.:\n\n   ```\n   [INFO] Current installed device driver is: 580.126.09\n   [INFO] Begin device health check\n   Device PCI Slot: 0000:00:03.0, Diagnosis result: 0000:00:03.0_1321122011772_0_0\n   * Power Cable Error Check - OK\n   * Device Driver Install Check - Failed\n   ... (one line per check item)\n   [INFO] Device health check completed\n   ```\n\n   > ⚠️ **Early-abort output**: When the GPU driver is NOT installed, the plugin exits EARLY (non-zero exit code, invocation status `Failed`/`ExitCodeNonzero`) with an output like below — this IS a valid finding and MUST be reported as `Device Driver Install Check – Failed` (IssueId `GuestOS.GPU.DriverNotInstalled`), NOT as a diagnosis failure:\n   >\n   > ```\n   > [ERROR] nvidia driver not installed\n   > [ERROR] Device driver not installed\n   > ```\n\n   Map failed check items to IssueIds per the table below:\n\n   | Check Item Name | Mapped IssueId | Diagnostic Description | Exception Handling Measures |\n   |-----------------|---------------|------------------------|----------------------------|\n   | Double Bit Error Check | GuestOS.GPU.MemoryEccCheckError | Detect GPU Double Bit Error conditions | Prompt user to restart instance based on error count |\n   | Info Rom Corrupted Check | GuestOS.GPU.InfoRomCorrupted | Detect GPU infoROM firmware information | O&M notification will be sent to user |\n   | eRDMA Incorrect Check | — (no mapped IssueId) | Detect GPU eRDMA network card status | O&M notification will be sent to user |\n   | Kernel Upgrade Check | GuestOS.GPU.DriverVersionMismatch | Detect driver anomalies caused by Kernel upgrades | User needs to uninstall and reinstall driver |\n   | Fabricmanager running Check | GuestOS.GPU.FabricmanagerCheck | Detect Fabricmanager component running status | User needs to install or start Fabricmanager component service |\n   | Power Cable Error Check | GuestOS.GPU.PowerCableError | Detect GPU power cable and power supply status | O&M notification will be sent to user |\n   | Device Lost Check | GuestOS.GPU.DeviceLost | Detect GPU card loss conditions | O&M notification will be sent to user |\n   | Device Physical Lost Check | GuestOS.GPU.DeviceLost | Detect GPU physical card loss conditions | O&M notification will be sent to user |\n   | Device Driver Install Check | GuestOS.GPU.DriverNotInstalled | Detect GPU driver installation status | User needs to install driver |\n   | Device Xid Error Check | GuestOS.GPU.NVXidError | Detect GPU Xid error anomalies | Prompt user to restart instance based on different XID errors |\n   | NVLink state Check | GuestOS.GPU.NVLinkError | Check GPU NVlink status | O&M notification will be sent to user |\n\n   **Special Reminder**: For check items whose handling measure is \"O&M notification will be sent to user\", append the reminder defined in the Output Description section.\n\n   If all check items are OK, the Cloud Assistant Diagnosis for this instance is considered normal.\n\n---\n\n#### Scheduled Diagnosis (Cloud Assistant Diagnosis ONLY)\n\n> **Scheduled diagnosis supports Cloud Assistant Diagnosis ONLY.** Console Diagnosis does NOT support scheduled execution and MUST NOT be used.\n>\n> Creates a Cloud Assistant command and a periodic schedule (`CreateCommand` + `InvokeCommand` with `--frequency`). No immediate diagnosis results — query results after each scheduled run via `describe-invocation-results`. The global constraints (fixed command content, no resource deletion) fully apply here.\n\n1. **Collect Parameters**\n   - `REGION_ID` (same validation as Prerequisites)\n   - Instance selection — ONE of: **explicit instance IDs** from the user; or **instance tag filtering** with tag Key and optionally Value (e.g., `gpu`, or `gpu=1`; no Value = match ALL values of the Key)\n   - Schedule time: user gives a time expression OR natural language (e.g., \"every day at 10:00\", \"every 2 hours\"). Convert it into a **6-field Cron expression** (`seconds minutes hours day month weekday`), e.g., every day at 10:00 → `0 0 10 * * ?`, every 6 hours → `0 0 */6 * * ?`\n   - Command name: `gpu_diagnosis_<date>` (`<date>` = today, `YYYYMMDD`); if the name exists, append a distinguishing suffix (e.g., `gpu_diagnosis_20260807_tag`)\n\n2. **Resolve Target Instances**\n   - Explicit instance IDs: verify existence and OS type via `describe-instances` (same as Prerequisites step 4); only Linux instances proceed\n   - Tag filtering: query matching instances first:\n     ```bash\n     aliyun ecs describe-instances \\\n       --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n       --biz-region-id '${REGION_ID}' \\\n       --region '${REGION_ID}' \\\n       --tag Key='${TAG_KEY}' Value='${TAG_VALUE}'   # omit Value= to match all values of the Key\n     ```\n     Extract the `InstanceId` list, filter to Linux instances only.\n   - > ⚠️ `invoke-command` does NOT support pure tag-based scheduling (it returns `MissingParam.InstanceId` when only `--tag` is passed). You MUST resolve the tag into explicit instance IDs BEFORE creating the schedule. Note that instances newly tagged later will NOT be automatically included; recreate the schedule if the instance set changes.\n\n3. **Create the Cloud Assistant Command**\n\n   The command content MUST be EXACTLY the same fixed Base64 literal used in Cloud Assistant Diagnosis step 1 (it decodes to the fixed single-line script). Use the literal DIRECTLY — do NOT re-encode or re-type the script at runtime:\n\n   ```bash\n   aliyun ecs create-command \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --name 'gpu_diagnosis_${DATE}' \\\n     --type RunShellScript \\\n     --command-content 'aWYgYWNzLXBsdWdpbi1tYW5hZ2VyIC0tbGlzdCAtLWxvY2FsIHwgZ3JlcCBBQ1MtRUNTLUdwdUNoZWNrID4gL2Rldi9udWxsIDI+JjE7IHRoZW4gYWNzLXBsdWdpbi1tYW5hZ2VyIC0tcmVtb3ZlIC0tcGx1Z2luIEFDUy1FQ1MtR3B1Q2hlY2s7IGZpOyBhY3MtcGx1Z2luLW1hbmFnZXIgLS1leGVjIC0tcGx1Z2luIEFDUy1FQ1MtR3B1Q2hlY2s='\n   ```\n\n   > ⚠️ `create-command` has NO `--content-encoding` flag — the API ALWAYS expects Base64 here. Before submitting, verify with `echo '<literal>' | base64 -d` that it decodes EXACTLY to the Cloud Assistant Diagnosis step 1 fixed single-line script. Pass the literal to `--command-content` directly; do NOT encode it again at runtime.\n\n   Extract `CommandId` from the response.\n\n4. **Create the Cron Schedule**\n\n   ```bash\n   aliyun ecs invoke-command \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --command-id '${COMMAND_ID}' \\\n     --instance-id '${INSTANCE_ID_1}' \\\n     --instance-id '${INSTANCE_ID_2}' \\\n     --frequency '${CRON_EXPRESSION}' \\\n     --timeout 180\n   ```\n\n   > ⚠️ Notes:\n   > - Multiple instances: repeat the `--instance-id` flag once per instance. Do NOT pass multiple IDs space-separated in a single flag (causes `InvalidInstance.NotFound`).\n   > - `--frequency` accepts the 6-field Cron expression (clock-based scheduling). The `--timed` parameter is deprecated — do NOT use it.\n\n   Extract `InvokeId` from the response.\n\n5. **Verify the Scheduled Task**\n\n   ```bash\n   aliyun ecs describe-invocations \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --invoke-id '${INVOKE_ID}'\n   ```\n\n   Confirm: `RepeatMode = Period`, `Frequency` matches the Cron expression, and each instance's `InvocationStatus = Scheduled`.\n\n6. **Report to the User**\n   - Output: command name, `CommandId`, `InvokeId`, Cron expression with human-readable schedule description, covered instance list\n   - Provide the result-query command for future runs: `aliyun ecs describe-invocation-results --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} --biz-region-id '${REGION_ID}' --region '${REGION_ID}' --invoke-id '${INVOKE_ID}'` (`Output` is Base64-encoded; decode and interpret per Cloud Assistant Diagnosis step 3)\n   - Remind: instances run against their status at trigger time; a stopped instance fails that run without affecting future schedules\n   - Do NOT execute any deletion commands (see Execution Constraints); if the user later asks to delete the task, stop it via `stop-invocation` and guide them to delete it manually in the console (see Edge Case Handling)\n\n### Output Description\n\nOutput results **by instance dimension**, with **abnormal instances listed first**, normal instances are NOT listed individually.\n\n**Output rules:**\n- **State the diagnosis conclusion DIRECTLY** — do NOT output any \"Diagnosis Complete!\" style banner or preamble\n- Use the method names **Console Diagnosis** and **Cloud Assistant Diagnosis** in the output; do NOT use \"Channel A\" / \"Channel B\" / \"dual-channel\" terminology, and do NOT display Report ID / Invoke ID / PCI slot details\n- The two methods are PEER-LEVEL, and output is organized BY METHOD DIMENSION (grouped by method): a **Console Diagnosis** section listing ITS anomaly items, followed by a **Cloud Assistant Diagnosis** section listing ITS anomaly items. Console Diagnosis findings are titled by their IssueId (e.g., `GuestOS.GPU.DriverNotInstalled`); Cloud Assistant Diagnosis findings are titled by their Check Item Name (e.g., `Device Driver Install Check - Failed`). Do NOT title a Cloud Assistant finding with a Console IssueId, and vice versa; if no anomalies exist, the conclusion states the instance/instances are normal\n- **Every anomaly item belongs to EXACTLY ONE method section** — NEVER merge findings of the two methods into one item, and NEVER write \"Detection method: Console Diagnosis + Cloud Assistant Diagnosis\". Findings from EITHER method MUST NEVER be dropped — see the \"No findings may be dropped, NEVER merge across methods\" constraint\n- **Grouped by method**: each method section has a header like `Console Diagnosis (N anomalies)` / `Cloud Assistant Diagnosis (N anomalies)`, and numbers its OWN items ([1], [2], ...); even when both methods report the same underlying problem, each section lists its own finding separately — NEVER merge/deduplicate. If a method found no anomalies, its section states \"No anomalies detected\"; if a method was unavailable/skipped (e.g., instance not running), its section states the reason\n- Only list instances that have anomalies found\n- Single-instance mode: method-grouped sections as above\n- Batch mode: output Cloud Assistant Diagnosis results only\n- At the end, show a summary line indicating how many instances are normal\n- If all instances are normal, only show the conclusion/summary\n- **Output language**: the format examples below are shown in English; render the SAME structure and labels in the user's conversation language\n- **Driver not installed**: When the anomaly is `GuestOS.GPU.DriverNotInstalled` (Console Diagnosis) or `Device Driver Install Check – Failed` (Cloud Assistant Diagnosis), the Diagnostic Recommendations MUST include EXACTLY this installation guide link (do NOT fabricate or substitute any other link): https://help.aliyun.com/zh/egs/install-a-gpu-driver-on-a-gpu-accelerated-compute-optimized-linux-instance . The recommendation MUST STOP at providing this official documentation link — do NOT ask the user whether they want you to install the driver, do NOT offer to install it on their behalf, and do NOT perform or attempt any driver installation on the instance\n\n**Single-instance output format:**\n\n```\nDiagnosis conclusion: instance i-bp1xxxxxxxxx (cn-shanghai) — 2 anomalies found (1 from Console Diagnosis, 1 from Cloud Assistant Diagnosis)\n\nConsole Diagnosis (1 anomaly):\n[1] GuestOS.GPU.DriverNotInstalled\n    Severity: Warn\n    Description: Detect GPU driver installation status\n    Action: User needs to install driver\n\nCloud Assistant Diagnosis (1 anomaly):\n[1] Device Driver Install Check - Failed\n    Description: Detect GPU driver installation status\n    Action: User needs to install driver\n\nRecommendations:\n- Install the matching version of the NVIDIA GPU driver\n- Installation guide: https://help.aliyun.com/zh/egs/install-a-gpu-driver-on-a-gpu-accelerated-compute-optimized-linux-instance\n```\n\n**Batch output format (Cloud Assistant Diagnosis only):**\n\n```\nDiagnosis conclusion: cn-shanghai — 5 instances in total: 2 abnormal, 3 normal\n\n========== ABNORMAL INSTANCES (2) ==========\n\n--- Instance: i-bp1aaaaaaaaaa ---\n[1] GuestOS.GPU.DriverNotInstalled — Detect GPU driver installation status. Action: install driver\n[2] GuestOS.GPU.NVXidError — Detect GPU Xid error anomalies. Action: restart instance based on the XID error\nRecommendations: Install NVIDIA driver; Restart instance to clear Xid errors\n\n... (repeat per abnormal instance)\n\n========== NORMAL INSTANCES (3) ==========\ni-bp3ccccccccccc, i-bp4ddddddddddd, i-bp5eeeeeeeeeee — No anomalies detected\n```\n\n**Special Reminder**: When the exception handling measure is \"O&M notification will be sent to user\", append the following reminder:\n```\n⚠️ Important Reminder:\n- Alibaba Cloud will send you O&M event notifications\n- Please go to the ECS console to view event details\n- Pay attention to whether you receive O&M events and handle them as required\n```\n\n### Edge Case Handling\n\n- **Instance does not exist**: CLI will return an error, capture and inform the user that the instance ID may be incorrect\n- **Region error**: Prompt user to confirm the region where the instance is located\n- **Non-GPU specification**: If the instance is not a GPU specification, diagnosis may have no results, prompt user to confirm instance type\n- **Insufficient permissions**: If permission error is returned, prompt user to check AccessKey permissions\n- **Network timeout**: Set command execution timeout (recommended 30 seconds), retry after timeout or prompt user to check network\n- **Cloud Assistant not available**: If RunCommand returns an error indicating the cloud assistant is not installed or the instance is not running, inform the user: \"Cloud Assistant Diagnosis is unavailable on instance ${INSTANCE_ID}. Please confirm the instance is in the Running state and the Cloud Assistant Agent is installed. Cloud Assistant Diagnosis for this instance has been skipped.\"\n- **Partial failure in batch**: If some instances fail in Cloud Assistant Diagnosis (e.g., cloud assistant unavailable), continue with the remaining instances and report failures separately\n- **Scheduled: pure tag scheduling rejected**: If `invoke-command` returns `MissingParam.InstanceId` when only `--tag` is passed, resolve the tag into instance IDs via `describe-instances` first, then create the schedule with explicit `--instance-id` flags\n- **Scheduled: command name conflict**: If `create-command` fails due to a duplicate name, append a distinguishing suffix to `gpu_diagnosis_<date>` and retry\n- **Scheduled: user requests deletion**: If the user asks to delete a scheduled task or command, do NOT delete any resource yourself per the \"Resource deletion is STRICTLY FORBIDDEN\" constraint. Instead: 1) execute `stop-invocation` (with the task's `InvokeId`) to stop future scheduled runs — stopping is permitted and is the agent's way to halt the task; 2) tell the user that deletion must be done manually in the console: ECS console → Cloud Assistant → Command Execution Results → Scheduled Executions, locate the scheduled task and delete it there\n\n### Example Workflow\n\n**Single Instance:**\n\n```\nUser: Help me diagnose this GPU server i-bp1xxxxxxxxx\n\nAgent:\n1. Check CLI is installed\n2. Ask for region (user did not provide)\n3. User replies: cn-shanghai\n4. Check instance OS type is Linux\n5. [Console Diagnosis] Execute CreateDiagnosticReport, get ReportId: dr-xxxxxxxx\n   [Cloud Assistant Diagnosis] Execute RunCommand, get InvokeId: t-xxxxxxxx (started simultaneously)\n6. Poll both DescribeDiagnosticReports and DescribeInvocationResults\n7. Cloud Assistant Diagnosis finishes first — parse Base64 output (even if status is Failed/ExitCodeNonzero, decode Output)\n8. Console Diagnosis finishes — parse Issues\n9. Per Output Description, output BY METHOD DIMENSION — a Console Diagnosis section listing its anomaly items, then a Cloud Assistant Diagnosis section listing its anomaly items (NEVER merge the two methods' anomalies into one item), output to user\n```\n\n**Scheduled Diagnosis (tag filtering):**\n\n```\nUser: Create a scheduled GPU diagnosis task for instances tagged gpu=1, running every day at 10:00\n\nAgent:\n1. Check CLI, validate region; resolve tag gpu=1 via describe-instances → Linux instances matched\n2. Convert \"every day at 10:00\" to Cron: 0 0 10 * * ?\n3. CreateCommand with the fixed script (Base64 literal, same as Cloud Assistant Diagnosis step 1), name gpu_diagnosis_<today>\n4. InvokeCommand with repeated --instance-id flags + --frequency '0 0 10 * * ?'\n5. Verify via describe-invocations (RepeatMode=Period, Scheduled); report IDs + result-query command\n6. NEVER delete the command or task yourself; if later asked to delete, run stop-invocation and direct the user to delete it in the ECS console (Cloud Assistant → Command Execution Results → Scheduled Executions)\n```\n\n**Batch Diagnosis:**\n\n```\nUser: Batch diagnose these GPU instances: i-bp1aaa, i-bp2bbb, i-bp3ccc, region cn-shanghai\n\nAgent:\n1. Check CLI; validate all instance IDs and region\n2. Query all instances' OS type in one call, filter to Linux only\n3. [Cloud Assistant Diagnosis ONLY] One RunCommand call with all instance IDs → 1 InvokeId\n4. Poll DescribeInvocationResults; output each instance's result as soon as it is ready\n5. Final summary: list normal instance IDs\n```\n\nFile v0.0.2:_meta.json\n\n{\n  \"ownerId\": \"kn74p5w8ywv6prh40g0s82gmqh83nw54\",\n  \"slug\": \"alibabacloud-ecs-gpu-diagnosis\",\n  \"version\": \"0.0.2\",\n  \"publishedAt\": 1786946934709\n}\n\nFile v0.0.2:references/cli-installation.md\n\n# Alibaba Cloud CLI Installation Guide\n\n## macOS\n\n**Homebrew (Recommended):**\n\n```bash\nbrew install aliyun-cli\n```\n\n**Manual Installation:**\n\n```bash\n# Intel chip\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# Apple Silicon (M1/M2/M3/M4)\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Linux\n\n```bash\n# x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Windows\n\n1. Download: https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\n2. Extract to any directory (e.g., `C:\\aliyun-cli`)\n3. The directory MUST be added to the system PATH environment variable\n\n**PowerShell Installation:**\n\n```powershell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli\n[Environment]::SetEnvironmentVariable(\"Path\", $env:Path + \";C:\\aliyun-cli\", [System.EnvironmentVariableTarget]::Machine)\n```\n\n## Verify Installation\n\nAfter installation, the following command MUST be executed to confirm successful installation:\n\n```bash\naliyun version\n```\n\nA version number in the output indicates successful installation. The version MUST be >= 3.0.299 to support OAuth authentication.\n\nAfter installation, the CLI MUST be updated to the latest version:\n\n```bash\n# macOS (Homebrew)\nbrew upgrade aliyun-cli\n\n# macOS (manual) / Linux: re-download the latest version and overwrite\n# Intel Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Apple Silicon Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Windows PowerShell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli -Force\n```\n\nAfter updating, `aliyun version` MUST be executed again to confirm the version has been updated.\n\n## Common Issues\n\n**command not found:** You MUST verify that the directory containing `aliyun` has been added to PATH.\n\n```bash\nwhich aliyun    # macOS/Linux\nwhere aliyun    # Windows\n```\n\n## Reference\n\n- Official documentation: https://help.aliyun.com/zh/cli/install-cli\n\nFile v0.0.2:references/ram-policies.md\n\n# RAM Permission List\n\nRAM permissions required for this Skill execution:\n\n## Diagnostic Operation Permissions\n\n`ecs:CreateDiagnosticReport` — Create ECS instance diagnostic report\n\n`ecs:DescribeDiagnosticReports` — Query diagnostic report status and results\n\n## Instance Query Permissions (for prerequisite checks)\n\n`ecs:DescribeInstances` — Query ECS instance basic information, verify instance existence\n\n`ecs:DescribeRegions` — Query ECS supported regions\n\n## Cloud Assistant Permissions (for Cloud Assistant Diagnosis)\n\n`ecs:RunCommand` — Execute cloud assistant commands on ECS instances\n\n`ecs:DescribeInvocationResults` — Query cloud assistant command execution results\n\n## Scheduled Diagnosis Permissions (for scheduled/periodic diagnosis)\n\n`ecs:CreateCommand` — Create a cloud assistant command with the fixed GPU diagnosis script\n\n`ecs:InvokeCommand` — Create a periodic schedule (Cron `--frequency`) to run the command on target instances\n\n`ecs:DescribeInvocations` — Verify the scheduled task status (RepeatMode, Frequency, InvocationStatus)\n\n`ecs:StopInvocation` — Stop a scheduled task when the user explicitly requests stopping/deleting it (stopping only; deletion is done manually in the ECS console)\n\nFile v0.0.2:skill-card.md\n\n## Description:\n\nDiagnoses GPU device status, driver installation issues, and hardware failures on Alibaba Cloud ECS GPU instances using Console Diagnosis and Cloud Assistant diagnosis modes.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[sdk-team](https://clawhub.ai/user/sdk-team)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and cloud operations engineers use this skill to diagnose single-instance, batch, or scheduled GPU health checks for Linux Alibaba Cloud ECS GPU instances and interpret Console Diagnosis and Cloud Assistant results.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill gives agents high-impact Alibaba Cloud command authority for ECS diagnostic and Cloud Assistant operations.\n\nMitigation: Use a tightly scoped RAM identity with only the documented permissions and review commands before execution.\n\nRisk: Untrusted instance lists, tag values, or region inputs could direct diagnosis commands at unintended resources.\n\nMitigation: Use explicit instance lists when possible, validate region and instance IDs, and avoid untrusted tag or region inputs.\n\nRisk: CLI installation and plugin update steps depend on mutable external sources.\n\nMitigation: Verify Alibaba Cloud CLI and plugin sources independently before installation or update.\n\nRisk: Scheduled diagnosis creates persistent cloud automation.\n\nMitigation: Track scheduled invocations, use stop-invocation only when explicitly requested, and clean up scheduled resources manually in the console.\n\n## Reference(s):\n\n- [Alibaba Cloud CLI Installation Guide](references/cli-installation.md)\n- [RAM Permission List](references/ram-policies.md)\n- [Alibaba Cloud CLI official installation documentation](https://help.aliyun.com/zh/cli/install-cli)\n- [Alibaba Cloud GPU driver installation guide](https://help.aliyun.com/zh/egs/install-a-gpu-driver-on-a-gpu-accelerated-compute-optimized-linux-instance)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and diagnosis summaries]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Streams diagnosis results as each method or instance completes; findings are grouped by diagnosis method.]\n\n## Skill Version(s):\n\n0.0.2 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.0.1: 5 files, 7006 bytes\n\nFiles: references/cli-installation.md (3085b), references/ram-policies.md (465b), skill-card.md (2679b), SKILL.md (10282b), _meta.json (149b)\n\nFile v0.0.1:SKILL.md\n\n---\nname: alibabacloud-ecs-gpu-diagnosis\ndescription: >\n  Diagnose Alibaba Cloud ECS GPU instances to detect GPU device status, driver issues, and hardware failures.\n  Use this Skill when users report GPU instance anomalies, deep learning task failures, GPU not visible, or when troubleshooting GPU hardware issues.\n  Supports automatic Alibaba Cloud CLI installation, diagnosis report creation, and polling for diagnosis results.\nlicense: Apache-2.0\ncompatibility: >\n  Requires Alibaba Cloud CLI (aliyun) installed and configured with AccessKey.\n  Supported regions: cn-hangzhou, cn-shanghai, cn-beijing, cn-shenzhen, etc.\nmetadata:\n  domain: aiops\n  owner: ecs-team\n  contact: ecs-agent@alibaba-inc.com\n---\n\n## Usage Instructions\n\nInitiate diagnosis on a specified ECS GPU instance to detect GPU device status and output diagnosis results.\n\n## Execution Constraints\n\n- All steps MUST be executed in order; skipping steps is NOT permitted\n- Each step MUST be verified as successful before proceeding to the next\n- Inform the user of the current step being executed\n- If any step fails, user confirmation MUST be obtained before continuing\n\n### Prerequisites\n\n1. **Check Alibaba Cloud CLI Environment**\n   - Execute `which aliyun` or `aliyun --version` to check if CLI is installed\n   - If not installed, inform the user that Alibaba Cloud CLI needs to be installed and provide installation guidance from `references/cli-installation.md`:\n     - macOS: Homebrew installation or manual installation (Intel/Apple Silicon)\n     - Linux: Download installation package for corresponding architecture (x86_64/ARM64)\n     - Windows: Download installation package and configure PATH, or use PowerShell installation\n   - After installation, run `aliyun version` to confirm version >= 3.0.299\n   - MUST run `aliyun plugin update` to ensure local plugins are up-to-date.\n   - Confirm CLI is configured with AccessKey: `aliyun configure`\n   - At the **start** of the Core Workflow (before any CLI invocation):\n    - **[MUST] Enable AI-Mode** — AI-mode is required for Agent Skill execution.\n     ```bash\n     aliyun configure ai-mode enable\n     aliyun configure ai-mode set-user-agent --user-agent \"AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis\"\n     ```\n   - At the **end** of the workflow (after success or failure):\n    - **[MUST] Disable AI-Mode** — AI-mode must be disabled after skill execution completes.\n     ```bash\n     aliyun configure ai-mode disable\n     ```\n   - **Permission Reminder**: Remind the user that the current RAM user needs the permissions to execute GPU diagnosis from `references/ram-policies.md` :\n\n2. **Obtain Required Parameters**\n   - Check if `INSTANCE_ID` is provided (ECS instance ID, format MUST match this regular expression ^i-[a-z0-9]{20}$ )\n   - Check if `REGION_ID` is provided (region ID, like cn-shanghai)\n   - If either parameter is missing, ask the user:\n     - \"Please provide the ECS instance ID to diagnose (format: i-bp1xxxxx)\"\n     - \"Please provide the region ID where the instance is located (e.g., cn-shanghai, cn-hangzhou)\"\n\n3. **Validate Parameters**\n   - **Validate INSTANCE_ID format**: Check if `INSTANCE_ID` matches the regex pattern `^i-[a-z0-9]{20}$`\n     - If validation fails, inform the user: \"Invalid instance ID format. Instance ID must match the pattern ^i-[a-z0-9]{20}$\"\n   - **Validate REGION_ID**: Query available regions using describe-regions API to verify the region is valid:\n     ```bash\n     aliyun ecs describe-regions --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis\n     ```\n     - Extract the `Regions.Region[].RegionId` list from the response\n     - Check if the provided `REGION_ID` exists in the list\n     - If region is invalid, inform the user: \"Invalid region ID. Please provide a valid region ID from the available regions list.\"\n\n4. **Check Instance Operating System Type**\n   - Before creating a diagnosis report, query instance information to confirm the OS type:\n     ```bash\n     aliyun ecs describe-instances  --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis --RegionId ${REGION_ID} --InstanceIds '[\"${INSTANCE_ID}\"]'\n     ```\n   - Extract the `Instances.Instance[0].OSType` field from the response\n   - **If `OSType` is \"linux\"**: Continue with the subsequent diagnosis process\n   - **If `OSType` is not \"linux\"**: Notify the user and terminate the process:\n     ```\n     The current instance ${INSTANCE_ID} has operating system ${OSType}.\n     This Skill currently only supports Linux operating system instances, other operating systems are not supported.\n     No further diagnosis process is needed.\n     ```\n\n### Execute Diagnosis\n\n1. **Create Diagnostic Report**\n   MUST Use the following command to initiate GPU diagnosis. Using alternatives such as RunCommand, nvidia-smi, DescribeInvocationResults, etc., for GPU diagnostics is prohibited.\n\n   ```bash\n   aliyun ecs create-diagnostic-report \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis \\\n     --RegionId '${REGION_ID}' \\\n     --ResourceId '${INSTANCE_ID}' \\\n     --MetricSetId 'dms-instanceGPUdevice' \\\n     --output cols=ReportId\n   ```\n\n   Extract `ReportId` from the output and save it for subsequent queries.\n\n2. **Poll Diagnostic Results**\n\n   MUST Use the following command to query the diagnosis report status:\n\n   ```bash\n   aliyun ecs describe-diagnostic-reports \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis \\\n     --RegionId '${REGION_ID}' \\\n     --ReportIds.1 '${REPORT_ID}'\n   ```\n\n   Handle based on the returned `Status` field:\n\n   - **Status = \"Finished\"**: Diagnosis complete, parse the `Issues` field content\n     - If `Issues` is empty or does not exist, report \"GPU diagnosis normal, no anomalies detected\"\n     - If `Issues` contains content, extract each Issue's `IssueId`, `MetricId`, `Severity`, and `MetricCategory`, and output diagnosis results and recommended actions according to the IssueId mapping table below\n   - **Status = \"InProgress\"**: Diagnosis in progress, wait 5 seconds before querying again\n   - **Status = \"Failed\"**: Diagnosis failed, report the failure status to the user\n\n   Set timeout mechanism: poll up to 60 times (approximately 5 minutes), if still not complete, prompt the user to query manually later.\n\n### Output Description\n\nAfter diagnosis is complete, the output should include:\n- Instance ID and region\n- Diagnostic report ID\n- GPU device status summary\n- Discovered Issues (if any)\n- Recommended remediation measures (inferred from Issues content)\n\n### Diagnostic Result Analysis\n\nThe `Issues` returned in the diagnosis report is an array, where each Issue contains `IssueId`, `MetricId`, `Severity`, and `MetricCategory` fields. Output diagnosis description and handling measures according to the IssueId mapping table below:\n\n| IssueId | Diagnostic Description | Exception Handling Measures |\n|---------|------------------------|----------------------------|\n| GuestOS.GPU.MemoryEccCheckError | Detect GPU Double Bit Error conditions | Prompt user to restart instance based on error count |\n| GuestOS.GPU.InfoRomCorrupted | Detect GPU infoROM firmware information | O&M notification will be sent to user |\n| GuestOS.GPU.DriverVersionMismatch | Detect driver anomalies caused by Kernel upgrades | User needs to uninstall and reinstall driver |\n| GuestOS.GPU.FabricmanagerCheck | Detect Fabricmanager component running status | User needs to install or start Fabricmanager component service |\n| GuestOS.GPU.PowerCableError | Detect GPU power cable and power supply status | O&M notification will be sent to user |\n| GuestOS.GPU.DeviceLost | Detect GPU card loss conditions | O&M notification will be sent to user |\n| GuestOS.GPU.DriverNotInstalled | Detect GPU driver installation status | User needs to install driver |\n| GuestOS.GPU.NVXidError | Detect GPU Xid error anomalies | Prompt user to restart instance based on different XID errors |\n| GuestOS.GPU.RmInitAdapterError | Detect GPU card initialization anomalies, manifested as driver card loss | O&M notification will be sent to user |\n| GuestOS.GPU.NVLinkError | Check GPU NVlink status | O&M notification will be sent to user |\n\n**Output Format Example**:\n\n```\nDiagnosis Complete! Instance: i-bp1xxxxxxxxx (cn-shanghai)\nReport ID: dr-xxxxxxxx\n\n1 anomaly found:\n\n[1] GuestOS.GPU.DriverNotInstalled\n    Severity: Warn\n    Diagnostic Description: Detect GPU driver installation status\n    Handling Measures: User needs to install driver\n\nDiagnostic Recommendations:\n- Please install the corresponding version of NVIDIA GPU driver\n- Installation Guide: https://help.aliyun.com/document_detail/108460.html\n```\n\n**Special Reminder**: When the exception handling measure is \"O&M notification will be sent to user\", append the following reminder to the output:\n```\n⚠️ Important Reminder:\n- Alibaba Cloud will send you O&M event notifications\n- Please go to the ECS console to view event details\n- Pay attention to whether you receive O&M events and handle them as required\n```\n\nIf `Issues` is an empty array or does not exist, output:\n```\nDiagnosis Complete! Instance: i-bp1xxxxxxxxx (cn-shanghai)\nReport ID: dr-xxxxxxxx\n\nGPU diagnosis normal, no anomalies detected.\n```\n\n### Edge Case Handling\n\n- **Instance does not exist**: CLI will return an error, capture and inform the user that the instance ID may be incorrect\n- **Region error**: Prompt user to confirm the region where the instance is located\n- **Non-GPU specification**: If the instance is not a GPU specification, diagnosis may have no results, prompt user to confirm instance type\n- **Insufficient permissions**: If permission error is returned, prompt user to check AccessKey permissions\n- **Network timeout**: Set command execution timeout (recommended 30 seconds), retry after timeout or prompt user to check network\n\n### Example Workflow\n\n```\nUser: Help me diagnose this GPU server i-bp1xxxxxxxxx\n\nAgent:\n1. Check CLI is installed\n2. Ask for region (user did not provide)\n3. User replies: cn-shanghai\n4. Check instance OS type is Linux\n5. Execute CreateDiagnosticReport, get ReportId: dr-xxxxxxxx\n6. Poll DescribeDiagnosticReports\n7. Status=InProgress, wait 5 seconds...\n8. Query again, Status=Finished\n9. Output Issues content to user\n```\n\nFile v0.0.1:_meta.json\n\n{\n  \"ownerId\": \"kn74p5w8ywv6prh40g0s82gmqh83nw54\",\n  \"slug\": \"alibabacloud-ecs-gpu-diagnosis\",\n  \"version\": \"0.0.1\",\n  \"publishedAt\": 1782358414297\n}\n\nFile v0.0.1:references/cli-installation.md\n\n# Alibaba Cloud CLI Installation Guide\n\n## macOS\n\n**Homebrew (Recommended):**\n\n```bash\nbrew install aliyun-cli\n```\n\n**Manual Installation:**\n\n```bash\n# Intel chip\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# Apple Silicon (M1/M2/M3/M4)\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Linux\n\n```bash\n# x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Windows\n\n1. Download: https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\n2. Extract to any directory (e.g., `C:\\aliyun-cli`)\n3. The directory MUST be added to the system PATH environment variable\n\n**PowerShell Installation:**\n\n```powershell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli\n[Environment]::SetEnvironmentVariable(\"Path\", $env:Path + \";C:\\aliyun-cli\", [System.EnvironmentVariableTarget]::Machine)\n```\n\n## Verify Installation\n\nAfter installation, the following command MUST be executed to confirm successful installation:\n\n```bash\naliyun version\n```\n\nA version number in the output indicates successful installation. The version MUST be >= 3.0.299 to support OAuth authentication.\n\nAfter installation, the CLI MUST be updated to the latest version:\n\n```bash\n# macOS (Homebrew)\nbrew upgrade aliyun-cli\n\n# macOS (manual) / Linux: re-download the latest version and overwrite\n# Intel Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Apple Silicon Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Windows PowerShell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli -Force\n```\n\nAfter updating, `aliyun version` MUST be executed again to confirm the version has been updated.\n\n## Common Issues\n\n**command not found:** You MUST verify that the directory containing `aliyun` has been added to PATH.\n\n```bash\nwhich aliyun    # macOS/Linux\nwhere aliyun    # Windows\n```\n\n## Reference\n\n- Official documentation: https://help.aliyun.com/zh/cli/install-cli\n\nFile v0.0.1:references/ram-policies.md\n\n# RAM Permission List\n\nRAM permissions required for this Skill execution:\n\n## Diagnostic Operation Permissions\n\n`ecs:CreateDiagnosticReport` — Create ECS instance diagnostic report\n\n`ecs:DescribeDiagnosticReports` — Query diagnostic report status and results\n\n## Instance Query Permissions (for prerequisite checks)\n\n`ecs:DescribeInstances` — Query ECS instance basic information, verify instance existence\n\n`ecs:DescribeRegions` - Query ECS supported regions\n\nFile v0.0.1:skill-card.md\n\n## Description: <br>\nDiagnose Alibaba Cloud ECS GPU instances to detect GPU device status, driver issues, and hardware failures. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[sdk-team](https://clawhub.ai/user/sdk-team) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nCloud operators, developers, and support engineers use this skill to validate Alibaba Cloud CLI readiness, create ECS GPU diagnostic reports, poll report status, and summarize detected GPU issues with recommended remediation steps. <br>\n\n### Deployment Geography for Use: <br>\nAlibaba Cloud ECS regions supported by the skill, including cn-hangzhou, cn-shanghai, cn-beijing, and cn-shenzhen. <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill uses Alibaba Cloud CLI credentials to query ECS instances and create diagnostic reports. <br>\nMitigation: Use least-privilege RAM credentials with only the required ECS diagnostic and instance query permissions, and confirm the target instance and region before execution. <br>\nRisk: CLI installation and update steps can modify a system-wide Alibaba Cloud CLI binary. <br>\nMitigation: Verify the CLI download source, review installation commands before running them, and prefer managed package installation where available. <br>\nRisk: AI-mode changes CLI configuration for agent execution. <br>\nMitigation: Enable AI-mode only for the diagnostic workflow and disable it after success or failure as the skill instructs. <br>\n\n\n## Reference(s): <br>\n- [Alibaba Cloud CLI Installation Guide](references/cli-installation.md) <br>\n- [RAM Permission List](references/ram-policies.md) <br>\n- [Alibaba Cloud CLI official installation documentation](https://help.aliyun.com/zh/cli/install-cli) <br>\n- [Alibaba Cloud GPU driver installation guide](https://help.aliyun.com/document_detail/108460.html) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, shell commands, guidance] <br>\n**Output Format:** [Markdown with Alibaba Cloud CLI commands, diagnostic report details, issue summaries, and remediation guidance] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Requires an ECS instance ID, region ID, configured Alibaba Cloud CLI credentials, and RAM permissions for ECS diagnostic and instance query APIs.] <br>\n\n## Skill Version(s): <br>\n0.0.1 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v0.0.1-beta.1: 5 files, 6905 bytes\n\nFiles: references/cli-installation.md (3085b), references/ram-policies.md (413b), skill-card.md (3038b), SKILL.md (9396b), _meta.json (156b)\n\nFile v0.0.1-beta.1:SKILL.md\n\n---\nname: alibabacloud-ecs-gpu-diagnosis\ndescription: >\n  Diagnose Alibaba Cloud ECS GPU instances to detect GPU device status, driver issues, and hardware failures.\n  Use this Skill when users report GPU instance anomalies, deep learning task failures, GPU not visible, or when troubleshooting GPU hardware issues.\n  Supports automatic Alibaba Cloud CLI installation, diagnosis report creation, and polling for diagnosis results.\nlicense: Apache-2.0\ncompatibility: >\n  Requires Alibaba Cloud CLI (aliyun) installed and configured with AccessKey.\n  Supported regions: cn-hangzhou, cn-shanghai, cn-beijing, cn-shenzhen, etc.\nmetadata:\n  domain: aiops\n  owner: ecs-team\n  contact: ecs-agent@alibaba-inc.com\n---\n\n## Usage Instructions\n\nInitiate diagnosis on a specified ECS GPU instance to detect GPU device status and output diagnosis results.\n\n## Execution Constraints\n\n- All steps MUST be executed in order; skipping steps is NOT permitted\n- Each step MUST be verified as successful before proceeding to the next\n- Inform the user of the current step being executed\n- If any step fails, user confirmation MUST be obtained before continuing\n\n### Prerequisites\n\n1. **Check Alibaba Cloud CLI Environment**\n   - Execute `which aliyun` or `aliyun --version` to check if CLI is installed\n   - If not installed, inform the user that Alibaba Cloud CLI needs to be installed and provide installation guidance from `references/cli-installation.md`:\n     - macOS: Homebrew installation or manual installation (Intel/Apple Silicon)\n     - Linux: Download installation package for corresponding architecture (x86_64/ARM64)\n     - Windows: Download installation package and configure PATH, or use PowerShell installation\n   - After installation, run `aliyun version` to confirm version >= 3.0.299\n   - Confirm CLI is configured with AccessKey: `aliyun configure`\n   - **Permission Reminder**: Remind the user that the current RAM user needs the permissions to execute GPU diagnosis from `references/ram-policies.md` :\n\n2. **Obtain Required Parameters**\n   - Check if `INSTANCE_ID` is provided (ECS instance ID, format MUST match this regular expression ^i-[a-z0-9]{20}$ )\n   - Check if `REGION_ID` is provided (region ID, like cn-shanghai)\n   - If either parameter is missing, ask the user:\n     - \"Please provide the ECS instance ID to diagnose (format: i-bp1xxxxx)\"\n     - \"Please provide the region ID where the instance is located (e.g., cn-shanghai, cn-hangzhou)\"\n\n3. **Validate Parameters**\n   - **Validate INSTANCE_ID format**: Check if `INSTANCE_ID` matches the regex pattern `^i-[a-z0-9]{20}$`\n     - If validation fails, inform the user: \"Invalid instance ID format. Instance ID must match the pattern ^i-[a-z0-9]{20}$\"\n   - **Validate REGION_ID**: Query available regions using DescribeRegions API to verify the region is valid:\n     ```bash\n     aliyun ecs DescribeRegions --user-agent AlibabaCloud-Agent-Skills\n     ```\n     - Extract the `Regions.Region[].RegionId` list from the response\n     - Check if the provided `REGION_ID` exists in the list\n     - If region is invalid, inform the user: \"Invalid region ID. Please provide a valid region ID from the available regions list.\"\n\n4. **Check Instance Operating System Type**\n   - Before creating a diagnosis report, query instance information to confirm the OS type:\n     ```bash\n     aliyun ecs DescribeInstances  --user-agent AlibabaCloud-Agent-Skills --RegionId ${REGION_ID} --InstanceIds '[\"${INSTANCE_ID}\"]'\n     ```\n   - Extract the `Instances.Instance[0].OSType` field from the response\n   - **If `OSType` is \"linux\"**: Continue with the subsequent diagnosis process\n   - **If `OSType` is not \"linux\"**: Notify the user and terminate the process:\n     ```\n     The current instance ${INSTANCE_ID} has operating system ${OSType}.\n     This Skill currently only supports Linux operating system instances, other operating systems are not supported.\n     No further diagnosis process is needed.\n     ```\n\n### Execute Diagnosis\n\n1. **Create Diagnostic Report**\n\n   Use the following command to initiate GPU diagnosis:\n\n   ```bash\n   aliyun ecs CreateDiagnosticReport \\\n     --user-agent AlibabaCloud-Agent-Skills \\\n     --RegionId '${REGION_ID}' \\\n     --ResourceId '${INSTANCE_ID}' \\\n     --MetricSetId 'dms-instanceGPUdevice' \\\n     --output cols=ReportId\n   ```\n\n   Extract `ReportId` from the output and save it for subsequent queries.\n\n2. **Poll Diagnostic Results**\n\n   Use the following command to query the diagnosis report status:\n\n   ```bash\n   aliyun ecs DescribeDiagnosticReports \\\n     --user-agent AlibabaCloud-Agent-Skills \\\n     --RegionId '${REGION_ID}' \\\n     --ReportIds.1 '${REPORT_ID}'\n   ```\n\n   Handle based on the returned `Status` field:\n\n   - **Status = \"Finished\"**: Diagnosis complete, parse the `Issues` field content\n     - If `Issues` is empty or does not exist, report \"GPU diagnosis normal, no anomalies detected\"\n     - If `Issues` contains content, extract each Issue's `IssueId`, `MetricId`, `Severity`, and `MetricCategory`, and output diagnosis results and recommended actions according to the IssueId mapping table below\n   - **Status = \"InProgress\"**: Diagnosis in progress, wait 5 seconds before querying again\n   - **Status = \"Failed\"**: Diagnosis failed, report the failure status to the user\n\n   Set timeout mechanism: poll up to 60 times (approximately 5 minutes), if still not complete, prompt the user to query manually later.\n\n### Output Description\n\nAfter diagnosis is complete, the output should include:\n- Instance ID and region\n- Diagnostic report ID\n- GPU device status summary\n- Discovered Issues (if any)\n- Recommended remediation measures (inferred from Issues content)\n\n### Diagnostic Result Analysis\n\nThe `Issues` returned in the diagnosis report is an array, where each Issue contains `IssueId`, `MetricId`, `Severity`, and `MetricCategory` fields. Output diagnosis description and handling measures according to the IssueId mapping table below:\n\n| IssueId | Diagnostic Description | Exception Handling Measures |\n|---------|------------------------|----------------------------|\n| GuestOS.GPU.MemoryEccCheckError | Detect GPU Double Bit Error conditions | Prompt user to restart instance based on error count |\n| GuestOS.GPU.InfoRomCorrupted | Detect GPU infoROM firmware information | O&M notification will be sent to user |\n| GuestOS.GPU.DriverVersionMismatch | Detect driver anomalies caused by Kernel upgrades | User needs to uninstall and reinstall driver |\n| GuestOS.GPU.FabricmanagerCheck | Detect Fabricmanager component running status | User needs to install or start Fabricmanager component service |\n| GuestOS.GPU.PowerCableError | Detect GPU power cable and power supply status | O&M notification will be sent to user |\n| GuestOS.GPU.DeviceLost | Detect GPU card loss conditions | O&M notification will be sent to user |\n| GuestOS.GPU.DriverNotInstalled | Detect GPU driver installation status | User needs to install driver |\n| GuestOS.GPU.NVXidError | Detect GPU Xid error anomalies | Prompt user to restart instance based on different XID errors |\n| GuestOS.GPU.RmInitAdapterError | Detect GPU card initialization anomalies, manifested as driver card loss | O&M notification will be sent to user |\n| GuestOS.GPU.NVLinkError | Check GPU NVlink status | O&M notification will be sent to user |\n\n**Output Format Example**:\n\n```\nDiagnosis Complete! Instance: i-bp1xxxxxxxxx (cn-shanghai)\nReport ID: dr-xxxxxxxx\n\n1 anomaly found:\n\n[1] GuestOS.GPU.DriverNotInstalled\n    Severity: Warn\n    Diagnostic Description: Detect GPU driver installation status\n    Handling Measures: User needs to install driver\n\nDiagnostic Recommendations:\n- Please install the corresponding version of NVIDIA GPU driver\n- Installation Guide: https://help.aliyun.com/document_detail/108460.html\n```\n\n**Special Reminder**: When the exception handling measure is \"O&M notification will be sent to user\", append the following reminder to the output:\n```\n⚠️ Important Reminder:\n- Alibaba Cloud will send you O&M event notifications\n- Please go to the ECS console to view event details\n- Pay attention to whether you receive O&M events and handle them as required\n```\n\nIf `Issues` is an empty array or does not exist, output:\n```\nDiagnosis Complete! Instance: i-bp1xxxxxxxxx (cn-shanghai)\nReport ID: dr-xxxxxxxx\n\nGPU diagnosis normal, no anomalies detected.\n```\n\n### Edge Case Handling\n\n- **Instance does not exist**: CLI will return an error, capture and inform the user that the instance ID may be incorrect\n- **Region error**: Prompt user to confirm the region where the instance is located\n- **Non-GPU specification**: If the instance is not a GPU specification, diagnosis may have no results, prompt user to confirm instance type\n- **Insufficient permissions**: If permission error is returned, prompt user to check AccessKey permissions\n- **Network timeout**: Set command execution timeout (recommended 30 seconds), retry after timeout or prompt user to check network\n\n### Example Workflow\n\n```\nUser: Help me diagnose this GPU server i-bp1xxxxxxxxx\n\nAgent:\n1. Check CLI is installed\n2. Ask for region (user did not provide)\n3. User replies: cn-shanghai\n4. Check instance OS type is Linux\n5. Execute CreateDiagnosticReport, get ReportId: dr-xxxxxxxx\n6. Poll DescribeDiagnosticReports\n7. Status=InProgress, wait 5 seconds...\n8. Query again, Status=Finished\n9. Output Issues content to user\n```\n\nFile v0.0.1-beta.1:_meta.json\n\n{\n  \"ownerId\": \"kn74p5w8ywv6prh40g0s82gmqh83nw54\",\n  \"slug\": \"alibabacloud-ecs-gpu-diagnosis\",\n  \"version\": \"0.0.1-beta.1\",\n  \"publishedAt\": 1776305538867\n}\n\nFile v0.0.1-beta.1:references/cli-installation.md\n\n# Alibaba Cloud CLI Installation Guide\n\n## macOS\n\n**Homebrew (Recommended):**\n\n```bash\nbrew install aliyun-cli\n```\n\n**Manual Installation:**\n\n```bash\n# Intel chip\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# Apple Silicon (M1/M2/M3/M4)\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Linux\n\n```bash\n# x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Windows\n\n1. Download: https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\n2. Extract to any directory (e.g., `C:\\aliyun-cli`)\n3. The directory MUST be added to the system PATH environment variable\n\n**PowerShell Installation:**\n\n```powershell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli\n[Environment]::SetEnvironmentVariable(\"Path\", $env:Path + \";C:\\aliyun-cli\", [System.EnvironmentVariableTarget]::Machine)\n```\n\n## Verify Installation\n\nAfter installation, the following command MUST be executed to confirm successful installation:\n\n```bash\naliyun version\n```\n\nA version number in the output indicates successful installation. The version MUST be >= 3.0.299 to support OAuth authentication.\n\nAfter installation, the CLI MUST be updated to the latest version:\n\n```bash\n# macOS (Homebrew)\nbrew upgrade aliyun-cli\n\n# macOS (manual) / Linux: re-download the latest version and overwrite\n# Intel Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Apple Silicon Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Windows PowerShell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli -Force\n```\n\nAfter updating, `aliyun version` MUST be executed again to confirm the version has been updated.\n\n## Common Issues\n\n**command not found:** You MUST verify that the directory containing `aliyun` has been added to PATH.\n\n```bash\nwhich aliyun    # macOS/Linux\nwhere aliyun    # Windows\n```\n\n## Reference\n\n- Official documentation: https://help.aliyun.com/zh/cli/install-cli\n\nFile v0.0.1-beta.1:references/ram-policies.md\n\n# RAM Permission List\n\nRAM permissions required for this Skill execution:\n\n## Diagnostic Operation Permissions\n\n`ecs:CreateDiagnosticReport` — Create ECS instance diagnostic report\n\n`ecs:DescribeDiagnosticReports` — Query diagnostic report status and results\n\n## Instance Query Permissions (for prerequisite checks)\n\n`ecs:DescribeInstances` — Query ECS instance basic information, verify instance existence\n\nFile v0.0.1-beta.1:skill-card.md\n\n## Description: <br>\nDiagnoses Alibaba Cloud ECS GPU instances to detect GPU device status, driver issues, and hardware failures, including guidance for CLI setup, diagnostic report creation, and result polling. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[sdk-team](https://clawhub.ai/user/sdk-team) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nCloud operations engineers, GPU platform teams, and developers use this skill to diagnose Alibaba Cloud ECS GPU instances when GPU devices are missing, drivers fail, deep learning workloads behave unexpectedly, or hardware-health issues are suspected. <br>\n\n### Deployment Geography for Use: <br>\nAlibaba Cloud ECS regions supported by the account and CLI workflow, with examples in the skill including cn-hangzhou, cn-shanghai, cn-beijing, and cn-shenzhen. <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The workflow requires Alibaba Cloud credentials and permissions that can affect cloud resources. <br>\nMitigation: Use a least-privilege RAM user or role with only the diagnostic and instance-query permissions needed for the target ECS instance. <br>\nRisk: Creating diagnostic reports against the wrong account, region, or instance can expose or act on unintended cloud resources. <br>\nMitigation: Confirm the Alibaba Cloud account, region, and instance ID before running CreateDiagnosticReport or polling diagnostic results. <br>\nRisk: CLI installation guidance includes downloaded binaries, sudo moves, PATH changes, and overwrite-style updates. <br>\nMitigation: Prefer official package-manager or vendor-documented installation, verify the download source, and avoid machine-wide changes unless the user explicitly accepts them. <br>\n\n\n## Reference(s): <br>\n- [ClawHub release page](https://clawhub.ai/sdk-team/alibabacloud-ecs-gpu-diagnosis) <br>\n- [Alibaba Cloud CLI Installation Guide](artifact/references/cli-installation.md) <br>\n- [RAM Permission List](artifact/references/ram-policies.md) <br>\n- [Alibaba Cloud CLI official installation documentation](https://help.aliyun.com/zh/cli/install-cli) <br>\n- [Alibaba Cloud GPU driver installation guide](https://help.aliyun.com/document_detail/108460.html) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown diagnostic summary with inline Alibaba Cloud CLI commands and remediation guidance] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Includes instance ID, region, diagnostic report ID, GPU status summary, discovered issues, and recommended remediation measures when available.] <br>\n\n## Skill Version(s): <br>\n0.0.1-beta.1 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>","readmeExcerpt":"Skill: alibabacloud-ecs-gpu-diagnosis Owner: sdk-team Summary: Diagnose GPU issues on Alibaba Cloud ECS GPU instances: GPU device status, driver issues, and GPU hardware failures. Use when users ask to check the GPU status of their GPU instances, detect whether the GPU device is visible, verify that the GPU driver is installed correctly, or troubleshoot GPU anomalies such as GPU not visible or deep learning task fail","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"aliyun ecs describe-regions --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} --region ${REGION_ID}"},{"language":"bash","snippet":"aliyun ecs describe-instances --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} --biz-region-id ${REGION_ID} --region ${REGION_ID} --instance-ids '[\"${INSTANCE_ID_1}\",\"${INSTANCE_ID_2}\",...]'"},{"language":"text","snippet":"--user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id}"},{"language":"bash","snippet":"aliyun ecs describe-instances --biz-region-id cn-hangzhou --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6"},{"language":"bash","snippet":"# Python SDK script\nSKILL_SESSION_ID={session-id} python3 scripts/deploy.py\n\n# Terraform\nSKILL_SESSION_ID={session-id} terraform apply"},{"language":"bash","snippet":"aliyun ecs create-diagnostic-report \\\n     --user-agent AlibabaCloud-Agent-Skills/alibabacloud-ecs-gpu-diagnosis/{session-id} \\\n     --biz-region-id '${REGION_ID}' \\\n     --region '${REGION_ID}' \\\n     --resource-id '${INSTANCE_ID}' \\\n     --metric-set-id 'dms-instanceGPUdevice' \\\n     --output cols=ReportId"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: alibabacloud-ecs-gpu-diagnosis\ndescription: >\n  Diagnose GPU issues on Alibaba Cloud ECS GPU instances: GPU device status, driver issues, and GPU hardware failures.\n  Use when users ask to check the GPU status of their GPU instances, detect whether the GPU device is visible, verify that the GPU driver is installed correctly, or troubleshoot GPU anomalies such as GPU not visible or deep learning task failures.\n  Run Console Diagnosis or Cloud Assistant Diagnosis (RunCommand) to detect GPU hardware failures, perform batch diagnosis of GPU servers, or create scheduled (periodic) diagnosis tasks via CreateCommand and InvokeCommand with Cron.\n  Single-instance diagnosis runs Console Diagnosis and Cloud Assistant Diagnosis in parallel; batch and scheduled diagnosis use Cloud Assistant Diagnosis only. Supports streaming output of diagnostic results.\n---\n\n## Usage Instructions\n\nDiagnose GPU device status, driver issues, and hardware failures on ECS instances using the following two diagnosis methods, **depending on the diagnosis mode**:\n- **Console Diagnosis**: `CreateDiagnosticReport` API, creates a diagnostic report and polls for results.\n- **Cloud Assistant Diagnosis**: `RunCommand` remotely executes the GPU health check plugin (`ACS-ECS-GpuCheck`) on the instance.\n\n**Mode-dependent method selection:**\n- **Single-instance diagnosis** (immediate): Console Diagnosis + Cloud Assistant Diagnosis **in parallel** (both methods launched simultaneously)\n- **Batch diagnosis** (immediate): **Cloud Assistant Diagnosis ONLY** — one `RunCommand` call with all instance IDs\n- **Scheduled diagnosis**: **Cloud Assistant Diagnosis ONLY** — `CreateCommand` + `InvokeCommand` with a Cron schedule (Console Diagnosis does NOT support scheduling). See the \"Scheduled Diagnosis (Cloud Assistant Diagnosis ONLY)\" section.\n\n## Execution Constraints\n\n- All steps MUST be executed in order; skipping steps is NOT permitted\n- Each step MUST be verified as successful before proceeding to the next\n- Inform the user of the current step being executed\n- If any step fails, user confirmation MUST be obtained before continuing\n- **Single-instance diagnosis**: Console Diagnosis and Cloud Assistant Diagnosis MUST execute **in parallel**, launched simultaneously. Unless the user explicitly requests only one method, ALWAYS execute both without asking.\n- **Batch diagnosis**: Execute **Cloud Assistant Diagnosis ONLY** via one batch `RunCommand` call. Do NOT create or poll Console diagnostic reports for batch instances.\n- **Scheduled diagnosis**: Execute **Cloud Assistant Diagnosis ONLY** via `CreateCommand` + `InvokeCommand` (Cron schedule). Do NOT use Console Diagnosis for scheduled tasks.\n- **Resource deletion is STRICTLY FORBIDDEN**: NEVER execute any deletion operation — including `delete-command`, deleting instances, tags, or any other cloud resources — even if the user asks; refuse and explain this constraint. `stop-invocation` (stop, not delete) is permitted ONLY when the user exp"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn74p5w8ywv6prh40g0s82gmqh83nw54\",\n  \"slug\": \"alibabacloud-ecs-gpu-diagnosis\",\n  \"version\": \"0.0.2\",\n  \"publishedAt\": 1786946934709\n}"},{"path":"references/cli-installation.md","content":"# Alibaba Cloud CLI Installation Guide\n\n## macOS\n\n**Homebrew (Recommended):**\n\n```bash\nbrew install aliyun-cli\n```\n\n**Manual Installation:**\n\n```bash\n# Intel chip\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# Apple Silicon (M1/M2/M3/M4)\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Linux\n\n```bash\n# x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz\nsudo mv aliyun /usr/local/bin/\n\n# ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz\nsudo mv aliyun /usr/local/bin/\n```\n\n## Windows\n\n1. Download: https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\n2. Extract to any directory (e.g., `C:\\aliyun-cli`)\n3. The directory MUST be added to the system PATH environment variable\n\n**PowerShell Installation:**\n\n```powershell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli\n[Environment]::SetEnvironmentVariable(\"Path\", $env:Path + \";C:\\aliyun-cli\", [System.EnvironmentVariableTarget]::Machine)\n```\n\n## Verify Installation\n\nAfter installation, the following command MUST be executed to confirm successful installation:\n\n```bash\naliyun version\n```\n\nA version number in the output indicates successful installation. The version MUST be >= 3.0.299 to support OAuth authentication.\n\nAfter installation, the CLI MUST be updated to the latest version:\n\n```bash\n# macOS (Homebrew)\nbrew upgrade aliyun-cli\n\n# macOS (manual) / Linux: re-download the latest version and overwrite\n# Intel Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-amd64.tgz\ntar -xzf aliyun-cli-macosx-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Apple Silicon Mac\nwget https://aliyuncli.alicdn.com/aliyun-cli-macosx-latest-arm64.tgz\ntar -xzf aliyun-cli-macosx-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux x86_64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-amd64.tgz\ntar -xzf aliyun-cli-linux-latest-amd64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Linux ARM64\nwget https://aliyuncli.alicdn.com/aliyun-cli-linux-latest-arm64.tgz\ntar -xzf aliyun-cli-linux-latest-arm64.tgz && sudo mv aliyun /usr/local/bin/\n\n# Windows PowerShell\nInvoke-WebRequest -Uri \"https://aliyuncli.alicdn.com/aliyun-cli-windows-latest-amd64.zip\" -OutFile \"aliyun-cli.zip\"\nExpand-Archive -Path aliyun-cli.zip -DestinationPath C:\\aliyun-cli -Force\n```\n\nAfter updating, `aliyun version` MUST be executed again to confirm the version has been updated.\n\n## Common Issues\n\n**command not found:** You MUST verify that the directory containing `aliyun` has been added to PATH.\n\n```bash\nwhich aliyun    # macOS/Linux\nwhere aliyun    # Windows\n```"},{"path":"references/ram-policies.md","content":"# RAM Permission List\n\nRAM permissions required for this Skill execution:\n\n## Diagnostic Operation Permissions\n\n`ecs:CreateDiagnosticReport` — Create ECS instance diagnostic report\n\n`ecs:DescribeDiagnosticReports` — Query diagnostic report status and results\n\n## Instance Query Permissions (for prerequisite checks)\n\n`ecs:DescribeInstances` — Query ECS instance basic information, verify instance existence\n\n`ecs:DescribeRegions` — Query ECS supported regions\n\n## Cloud Assistant Permissions (for Cloud Assistant Diagnosis)\n\n`ecs:RunCommand` — Execute cloud assistant commands on ECS instances\n\n`ecs:DescribeInvocationResults` — Query cloud assistant command execution results\n\n## Scheduled Diagnosis Permissions (for scheduled/periodic diagnosis)\n\n`ecs:CreateCommand` — Create a cloud assistant command with the fixed GPU diagnosis script\n\n`ecs:InvokeCommand` — Create a periodic schedule (Cron `--frequency`) to run the command on target instances\n\n`ecs:DescribeInvocations` — Verify the scheduled task status (RepeatMode, Frequency, InvocationStatus)\n\n`ecs:StopInvocation` — Stop a scheduled task when the user explicitly requests stopping/deleting it (stopping only; deletion is done manually in the ECS console)"},{"path":"skill-card.md","content":"## Description:\n\nDiagnoses GPU device status, driver installation issues, and hardware failures on Alibaba Cloud ECS GPU instances using Console Diagnosis and Cloud Assistant diagnosis modes.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[sdk-team](https://clawhub.ai/user/sdk-team)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and cloud operations engineers use this skill to diagnose single-instance, batch, or scheduled GPU health checks for Linux Alibaba Cloud ECS GPU instances and interpret Console Diagnosis and Cloud Assistant results.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill gives agents high-impact Alibaba Cloud command authority for ECS diagnostic and Cloud Assistant operations.\n\nMitigation: Use a tightly scoped RAM identity with only the documented permissions and review commands before execution.\n\nRisk: Untrusted instance lists, tag values, or region inputs could direct diagnosis commands at unintended resources.\n\nMitigation: Use explicit instance lists when possible, validate region and instance IDs, and avoid untrusted tag or region inputs.\n\nRisk: CLI installation and plugin update steps depend on mutable external sources.\n\nMitigation: Verify Alibaba Cloud CLI and plugin sources independently before installation or update.\n\nRisk: Scheduled diagnosis creates persistent cloud automation.\n\nMitigation: Track scheduled invocations, use stop-invocation only when explicitly requested, and clean up scheduled resources manually in the console.\n\n## Reference(s):\n\n- [Alibaba Cloud CLI Installation Guide](references/cli-installation.md)\n- [RAM Permission List](references/ram-policies.md)\n- [Alibaba Cloud CLI official installation documentation](https://help.aliyun.com/zh/cli/install-cli)\n- [Alibaba Cloud GPU driver installation guide](https://help.aliyun.com/zh/egs/install-a-gpu-driver-on-a-gpu-accelerated-compute-optimized-linux-instance)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and diagnosis summaries]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Streams diagnosis results as each method or instance completes; findings are grouped by diagnosis method.]\n\n## Skill Version(s):\n\n0.0.2 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1766,"uniquenessScore":38,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T12:10:15.904Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T12:10:15.904Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T15:17:14.780Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}