{"id":"e0e7c3ff-f1a8-402c-aa40-33d2e7a50933","entityType":"agent","slug":"clawhub-patmenciu-modelpilot","name":"Ollama Model Pilot","canonicalUrl":"https://www.xpersona.co/agent/clawhub-patmenciu-modelpilot","canonicalPath":"/agent/clawhub-patmenciu-modelpilot","generatedAt":"2026-10-11T07:40:11.858Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:31:17.124Z","emptyReason":null},"description":"Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-th...","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17b7ke0py60etvzaczrwdabh587gytb:modelpilot","sourceUrl":"https://clawhub.ai/patmenciu/modelpilot","homepage":"https://clawhub.ai/patmenciu/skills/modelpilot","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/patmenciu/modelpilot","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/patmenciu/skills/modelpilot","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Ollama Model Pilot technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:31:17.124Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:31:17.124Z","emptyReason":null},"stars":null,"forks":null,"downloads":1159,"packageName":null,"latestVersion":"1.5.0","tractionLabel":"1.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:31:17.055Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T04:31:17.124Z","lastCrawledAt":"2026-10-11T04:31:17.055Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T04:31:17.055Z","lastVerifiedAt":null,"highlights":[{"version":"1.5.0","createdAt":"2026-06-08T14:26:51.217Z","changelog":"**Modelpilot v1.5.0 Changelog** - Major refactor: SKILL.md now emphasizes a strict local-only safety boundary, two-round replacement rule, and repeatable benchmark protocol for Ollama models. - Added: Benchmark example files (`examples/models.example.json`, `examples/prompts.example.json`, `outputs/example_report.md`) and local scripts for running/reporting benchmarks (`scripts/modelpilot_benchmark.py`, `scripts/modelpilot_report.py`). - Added: Unit test for the new reporting workflow (`tests/test_modelpilot_report.py`). - Removed: Skill metadata files (`CHANGELOG.md`, `skill-card.md`) now replaced by clearer in-skill documentation. - New: Required response format for benchmark and decision results to ensure consistent reporting. - Now explicitly prohibits cloud API calls, model downloading, and any non-local actions without user approval.","fileCount":10,"zipByteSize":12227},{"version":"1.4.1","createdAt":"2026-05-29T06:43:25.998Z","changelog":"- Added explicit activation rules so agents know when to use this skill. - Made Real-Task Benchmark the default workflow for model testing, benchmarking, or comparison, unless the user asks for a simple smoke test. - Introduced a standard benchmark response format for consistent model-testing results. - Merged and removed legacy skill-card.md file.","fileCount":5,"zipByteSize":16348},{"version":"1.4.0","createdAt":"2026-05-28T03:59:13.569Z","changelog":"Ollama Model Pilot v1.4.0 Enhanced Real-Task Benchmark support, upgrading model testing from simple speed checks to workflow-oriented evaluation. Added recommended real-task benchmark categories: short QA, long summary, structured output, and role-based workflow tasks. Clarified that a model should not be promoted into a production workflow based on speed alone; output quality, format stability, and task fit should also be checked. Updated README.md and CHANGELOG.md to reflect the v1.4 naming update, maintenance direction, and Real-Task Benchmark improvements. This update improves documentation structure and model evaluation guidance. It does not introduce automatic deletion, automatic downloading, or destructive operations. Officially renamed the skill to Ollama Model Pilot, while keeping modelpilot as the package name / slug for compatibility. Added naming notes to clarify that this skill was previously published as Model Pilot and Ollama Lifecycle Manager, and will be maintained under Ollama Model Pilot going forward. 强化 Real-Task Benchmark / 真实任务基准测试，将模型测试从简单测速升级为面向真实工作流的评估流程。 新增真实任务测试维度：短问答、长文摘要、结构化输出、真实角色任务。 明确模型进入正式工作流前，不应只看速度，还应检查输出质量、格式稳定性和任务适配度。 更新 README.md 和 CHANGELOG.md，同步说明 v1.4 的命名调整、维护策略和真实任务 Benchmark 升级。 本次更新包含文档结构和模型测试方法改进，不涉及自动删除、自动下载或破坏性操作。 正式统一名称为 Ollama Model Pilot，保留 modelpilot 作为安装包名 / slug。 增加旧名说明，明确本 skill 早期曾使用 Model Pilot 和 Ollama Lifecycle Manager 名称，后续以 Ollama Model Pilot 作为主线维护。","fileCount":5,"zipByteSize":15819},{"version":"1.3.0","createdAt":"2026-05-27T10:12:34.936Z","changelog":"v1.3.0: Rebrand to Model Pilot — new name and descriptions highlighting safety-first copilot approach. Bilingual SKILL.md + SKILL.en.md.","fileCount":4,"zipByteSize":17922}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17b7ke0py60etvzaczrwdabh587gytb:modelpilot","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17b7ke0py60etvzaczrwdabh587gytb:modelpilot` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/patmenciu/modelpilot before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T07:40:11.855Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:31:17.124Z","emptyReason":null},"readme":"Skill: Ollama Model Pilot\n\nOwner: patmenciu\n\nSummary: Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-th...\n\nTags: latest:1.5.0\n\nVersion history:\n\nv1.5.0 | 2026-06-08T14:26:51.217Z | user\n\n**Modelpilot v1.5.0 Changelog**\n\n- Major refactor: SKILL.md now emphasizes a strict local-only safety boundary, two-round replacement rule, and repeatable benchmark protocol for Ollama models.\n- Added: Benchmark example files (`examples/models.example.json`, `examples/prompts.example.json`, `outputs/example_report.md`) and local scripts for running/reporting benchmarks (`scripts/modelpilot_benchmark.py`, `scripts/modelpilot_report.py`).\n- Added: Unit test for the new reporting workflow (`tests/test_modelpilot_report.py`).\n- Removed: Skill metadata files (`CHANGELOG.md`, `skill-card.md`) now replaced by clearer in-skill documentation.\n- New: Required response format for benchmark and decision results to ensure consistent reporting.\n- Now explicitly prohibits cloud API calls, model downloading, and any non-local actions without user approval.\n\nv1.4.1 | 2026-05-29T06:43:25.998Z | user\n\n- Added explicit activation rules so agents know when to use this skill.\n- Made Real-Task Benchmark the default workflow for model testing, benchmarking, or comparison, unless the user asks for a simple smoke test.\n- Introduced a standard benchmark response format for consistent model-testing results.\n- Merged and removed legacy skill-card.md file.\n\nv1.4.0 | 2026-05-28T03:59:13.569Z | user\n\nOllama Model Pilot v1.4.0\nEnhanced Real-Task Benchmark support, upgrading model testing from simple speed checks to workflow-oriented evaluation.\nAdded recommended real-task benchmark categories: short QA, long summary, structured output, and role-based workflow tasks.\nClarified that a model should not be promoted into a production workflow based on speed alone; output quality, format stability, and task fit should also be checked.\nUpdated README.md and CHANGELOG.md to reflect the v1.4 naming update, maintenance direction, and Real-Task Benchmark improvements.\nThis update improves documentation structure and model evaluation guidance. It does not introduce automatic deletion, automatic downloading, or destructive operations.\nOfficially renamed the skill to Ollama Model Pilot, while keeping modelpilot as the package name / slug for compatibility.\nAdded naming notes to clarify that this skill was previously published as Model Pilot and Ollama Lifecycle Manager, and will be maintained under Ollama Model Pilot going forward.\n\n强化 Real-Task Benchmark / 真实任务基准测试，将模型测试从简单测速升级为面向真实工作流的评估流程。\n新增真实任务测试维度：短问答、长文摘要、结构化输出、真实角色任务。\n明确模型进入正式工作流前，不应只看速度，还应检查输出质量、格式稳定性和任务适配度。\n更新 README.md 和 CHANGELOG.md，同步说明 v1.4 的命名调整、维护策略和真实任务 Benchmark 升级。\n本次更新包含文档结构和模型测试方法改进，不涉及自动删除、自动下载或破坏性操作。\n正式统一名称为 Ollama Model Pilot，保留 modelpilot 作为安装包名 / slug。\n增加旧名说明，明确本 skill 早期曾使用 Model Pilot 和 Ollama Lifecycle Manager 名称，后续以 Ollama Model Pilot 作为主线维护。\n\nv1.3.0 | 2026-05-27T10:12:34.936Z | user\n\nv1.3.0: Rebrand to Model Pilot — new name and descriptions highlighting safety-first copilot approach. Bilingual SKILL.md + SKILL.en.md.\n\nArchive index:\n\nArchive v1.5.0: 10 files, 12227 bytes\n\nFiles: examples/models.example.json (335b), examples/prompts.example.json (1878b), outputs/example_report.md (1026b), README.md (2555b), scripts/modelpilot_benchmark.py (5259b), scripts/modelpilot_report.py (6045b), skill-card.md (2593b), SKILL.md (4950b), tests/test_modelpilot_report.py (2236b), _meta.json (129b)\n\nFile v1.5.0:SKILL.md\n\n---\nname: modelpilot\ndescription: Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-think verification, and local-only safety boundaries. It applies to local LLM evaluation, model replacement decisions, benchmark reports, installed-model audits, and Ollama workflow hygiene. Do not use it for cloud model APIs, downloading models, installing dependencies, or sending local data outside the machine.\n---\n\n# ModelPilot\n\nModelPilot is a local-only protocol for testing, comparing, promoting, replacing,\nand cleaning up Ollama models. It is designed for real work decisions, not leaderboard\nclaims.\n\n## Safety Boundary\n\nAlways keep the workflow local unless the user explicitly authorizes otherwise.\n\n- Do not call cloud model APIs.\n- Do not upload files, prompts, logs, paths, configs, or benchmark outputs.\n- Do not download, pull, install, upgrade, or delete models without explicit user approval.\n- Do not use real private documents as benchmark samples unless the user explicitly names the file for this task.\n- Use fictional examples for tests, documentation, and demos.\n- Treat model cleanup as a workflow dependency audit, not a disk-space optimization task.\n\n## Trigger Conditions\n\nUse this skill when the user asks to:\n\n- test an Ollama model\n- compare local models\n- decide whether a new model can replace an existing model\n- verify no-think behavior\n- build a local model benchmark report\n- audit installed models before cleanup\n- choose local models for coding, writing, RAG, automation, or structured output\n\n## Test Levels\n\nClassify the task before running anything.\n\n1. Smoke Test\n   Confirm the model is installed, runnable, and responsive.\n\n2. Speed Benchmark\n   Measure startup time, generation time, output length, and failure rate.\n\n3. Real-Task Benchmark\n   Use task-like prompts that match the user's actual workflow. Prefer fixed prompt\n   sets so results are comparable across models.\n\n4. Promotion Test\n   Decide whether a model can replace an existing workflow model. A promotion test\n   requires two independent benchmark rounds.\n\n## Two-Round Replacement Rule\n\nDo not recommend replacing a working model after a single run.\n\n- Round 1 checks: runnable, speed, output format, obvious quality failures, no-think leakage.\n- Round 2 checks: same prompt set, same model, repeatability, quality consistency, failure modes.\n- A model is only replacement-ready when both rounds pass the required tasks.\n- Keep the previous model and configuration available for rollback.\n- If structured output, no-think behavior, or long-context handling is unstable, do not use the model in automation.\n\n## Fixed Prompt Set\n\nPrefer a stable prompt file with fictional data. Include at least:\n\n- short Chinese or English Q&A\n- long-document summary\n- structured JSON or Markdown output\n- real-role workflow simulation\n- no-think verification prompt\n\nThe benchmark prompt set should be reused across candidate models. Do not compare\nmodels using different tasks unless the report clearly says so.\n\n## No-Think Verification\n\nNever assume a model is no-think just because the model name contains `nothink`.\n\nCheck:\n\n- model output does not include `<think>`, `</think>`, reasoning traces, or hidden-analysis markers\n- CLI or API flags are actually accepted by the runtime\n- Modelfile-level instructions are treated as weak constraints, not proof\n- structured outputs remain clean when no-think is enabled\n\nIf no-think fails, the model may still be useful for manual work, but it should not\nbe promoted into automated workflows that require clean output.\n\n## Standard Workflow\n\n1. Identify the current model, candidate model, task type, and replacement target.\n2. Confirm whether the user wants smoke, speed, real-task, or promotion testing.\n3. Build or reuse a fixed prompt set with fictional data unless the user explicitly authorizes a real file.\n4. Run two benchmark rounds for replacement decisions.\n5. Review mechanical results: failures, duration, output length, format checks, no-think leakage.\n6. Review semantic quality manually for real-task tasks.\n7. Return a concise decision: keep, observe, replace, or not recommended.\n8. Include rollback advice when replacement is recommended.\n\n## Suggested Local Scripts\n\nUse scripts only when they are available in this skill folder and fit the task.\n\n- `scripts/modelpilot_benchmark.py`: run local Ollama benchmark rounds and write JSON results.\n- `scripts/modelpilot_report.py`: summarize benchmark JSON into a Markdown decision report.\n\nDo not run scripts that download models, install dependencies, or call remote APIs.\n\n## Required Response Format\n\nWhen reporting results, include:\n\n```markdown\n## ModelPilot Result\n\n### Scope\n-\n\n### Models Tested\n-\n\n### Test Rounds\n-\n\n### Key Findings\n-\n\n### No-Think Check\n-\n\n### Replacement Decision\n-\n\n### Risks and Limits\n-\n\n### Rollback Advice\n-\n```\n\nFile v1.5.0:README.md\n\n# ModelPilot\n\nModelPilot is a local-only skill for testing, comparing, promoting, replacing,\nand cleaning up Ollama models.\n\nIt is built around a simple rule: a model should not replace an existing workflow\nmodel after one good run. Run the same fixed prompt set twice, review both rounds,\nthen decide.\n\n## What It Helps With\n\n- Compare local Ollama models on real tasks\n- Verify whether a `nothink` model actually suppresses thinking traces\n- Decide whether a candidate model can replace a current model\n- Produce compact benchmark reports\n- Audit models before cleanup without deleting anything automatically\n\n## Directory Layout\n\n```text\nmodelpilot/\n  SKILL.md\n  README.md\n  scripts/\n  examples/\n  tests/\n  outputs/\n```\n\n## Safety Defaults\n\n- Local Ollama only\n- No cloud model APIs\n- No uploads\n- No model downloads\n- No dependency installation\n- No automatic model deletion\n- Fictional examples only\n\n## Basic Usage\n\nPrepare a fictional prompt set:\n\n```bash\npython scripts/modelpilot_benchmark.py \\\n  --models llama3.2:latest qwen3:latest \\\n  --prompts examples/prompts.example.json \\\n  --rounds 2 \\\n  --output outputs/benchmark_results.json\n```\n\nCreate a Markdown report:\n\n```bash\npython scripts/modelpilot_report.py \\\n  --input outputs/benchmark_results.json \\\n  --output outputs/benchmark_report.md\n```\n\nThe scripts use Python standard library only. The benchmark script calls the local\n`ollama` command and expects the models to already be installed.\n\n## Replacement Rule\n\nA model can be considered replacement-ready only after two independent rounds using\nthe same fixed prompt set.\n\nRound 1 checks:\n\n- the model runs\n- response speed is acceptable\n- output format is stable\n- no-think behavior does not leak reasoning text\n\nRound 2 checks:\n\n- the same tasks still pass\n- failure modes do not repeat\n- quality is consistent enough for the target workflow\n\nIf either round fails on structured output, no-think behavior, or the user's core\ntask, keep the existing model.\n\n## Example Decision Labels\n\n- `replace_ready`: both rounds pass and manual review confirms quality\n- `observe`: usable, but has minor instability or incomplete evidence\n- `candidate_only`: only one round is complete\n- `not_recommended`: repeated failures, output pollution, or unsafe automation fit\n\n## Limits\n\nModelPilot does not prove general intelligence or leaderboard quality. It helps make\nlocal workflow decisions based on fixed tasks, repeatability, and clean output.\n\nThe report script can detect mechanical issues, but semantic quality still needs a\nhuman review.\n\nFile v1.5.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fj0qnbect2fmp59jj0vx34587hvns\",\n  \"slug\": \"modelpilot\",\n  \"version\": \"1.5.0\",\n  \"publishedAt\": 1780928811217\n}\n\nFile v1.5.0:outputs/example_report.md\n\n# ModelPilot Benchmark Report\n\n- Version: 1.5.0\n- Generated at: 2026-06-08T00:00:00+00:00\n- Local only: true\n- Rounds requested: 2\n- Prompt count: 5\n\n## Summary\n\n| Model | Decision | Reason | Success | Format | Think leak | Avg seconds |\n| --- | --- | --- | ---: | ---: | ---: | ---: |\n| example-candidate-a:latest | replace_ready | Two rounds passed mechanical checks. Human semantic review is still required. | 10/10 | 10/10 | 0 | 3.42 |\n| example-candidate-b:nothink | not_recommended | 1 outputs show possible thinking leakage. | 10/10 | 10/10 | 1 | 3.10 |\n\n## Replacement Decision\n\n- `example-candidate-a:latest` can be considered for replacement after manual semantic review.\n- `example-candidate-b:nothink` should not be promoted into automated workflows until no-think output is clean.\n- Keep the previous model and config available for rollback.\n\n## Risks and Limits\n\n- This example uses fictional data.\n- Mechanical checks do not prove semantic quality.\n- Do not delete old models based only on a benchmark report.\n\nFile v1.5.0:skill-card.md\n\n## Description:\n\nModelPilot helps agents test, compare, promote, replace, and clean up local Ollama models using repeatable two-round benchmarks, no-think checks, and local-only safety boundaries.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[patmenciu](https://clawhub.ai/user/patmenciu)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and engineers use this skill to evaluate installed local Ollama models on repeatable task prompts, no-think behavior, structured output stability, and replacement readiness. It supports local workflow decisions without cloud APIs, uploads, model downloads, or automatic model deletion.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Benchmark reports can overstate replacement readiness because mechanical checks do not prove semantic quality.\n\nMitigation: Require human semantic review and two clean benchmark rounds before replacing a workflow model; keep the previous model and configuration available for rollback.\n\nRisk: Prompts, logs, or benchmark outputs may contain sensitive local data if real files are used.\n\nMitigation: Use fictional prompt sets by default, only use user-named files for the task, and keep benchmark outputs local.\n\nRisk: No-think model names or instructions may not guarantee clean output.\n\nMitigation: Check outputs for thinking traces and avoid promoting a model into automation when leakage appears.\n\nRisk: Untrusted or edited benchmark JSON can produce misleading reports.\n\nMitigation: Review benchmark inputs and reports manually, and rerun benchmarks from trusted model and prompt configurations when results affect replacement decisions.\n\n## Reference(s):\n\n- [ClawHub Skill Page](https://clawhub.ai/patmenciu/skills/modelpilot)\n- [README](README.md)\n- [Example Benchmark Report](outputs/example_report.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with optional shell commands, JSON benchmark results, and Markdown reports]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Local-only Ollama workflow; benchmark scripts use installed models and fictional or explicitly selected prompt files.]\n\n## Skill Version(s):\n\n1.5.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.5.0:examples/models.example.json\n\n{\n  \"version\": \"1.5.0\",\n  \"current_model\": \"example-current:latest\",\n  \"candidate_models\": [\n    \"example-candidate-a:latest\",\n    \"example-candidate-b:nothink\"\n  ],\n  \"target_workflow\": \"fictional local document summary workflow\",\n  \"notes\": \"Use installed local Ollama models only. Do not pull or download models from this file.\"\n}\n\nFile v1.5.0:examples/prompts.example.json\n\n{\n  \"version\": \"1.5.0\",\n  \"description\": \"Fictional benchmark prompts for local Ollama model testing.\",\n  \"prompts\": [\n    {\n      \"id\": \"short_qa_001\",\n      \"title\": \"Short factual answer\",\n      \"category\": \"smoke\",\n      \"expected_format\": \"plain_text\",\n      \"prompt\": \"Answer in one short paragraph: What is the difference between a checklist and a test plan?\"\n    },\n    {\n      \"id\": \"summary_001\",\n      \"title\": \"Fictional project summary\",\n      \"category\": \"real_task\",\n      \"expected_format\": \"markdown_bullets\",\n      \"prompt\": \"Summarize the following fictional project note into 4 bullet points. Note: The Demo Archive team tested a local search prototype with 120 synthetic documents. The prototype found titles quickly but sometimes missed date filters. The team wants a small follow-up test before replacing the old tool. No customer data was used.\"\n    },\n    {\n      \"id\": \"structured_json_001\",\n      \"title\": \"Structured JSON output\",\n      \"category\": \"automation\",\n      \"expected_format\": \"json\",\n      \"prompt\": \"Return valid JSON only. Extract fields from this fictional task: Project Alpha needs a local-only model test by Friday. Owner is Demo User. Risk is unstable JSON output. Required keys: project, owner, due, risk.\"\n    },\n    {\n      \"id\": \"role_task_001\",\n      \"title\": \"Workflow recommendation\",\n      \"category\": \"real_task\",\n      \"expected_format\": \"markdown_sections\",\n      \"prompt\": \"You are helping evaluate a local model for a fictional note-cleanup workflow. Give a concise recommendation with sections: Decision, Reason, Risk, Next Step.\"\n    },\n    {\n      \"id\": \"nothink_001\",\n      \"title\": \"No-think leakage check\",\n      \"category\": \"nothink\",\n      \"expected_format\": \"plain_text\",\n      \"prompt\": \"Answer with the final answer only, no reasoning trace: Should a model be promoted after one benchmark run?\"\n    }\n  ]\n}\n\nArchive v1.4.1: 5 files, 16348 bytes\n\nFiles: CHANGELOG.md (1505b), README.md (1435b), skill-card.md (2492b), SKILL.md (32696b), _meta.json (129b)\n\nFile v1.4.1:SKILL.md\n\n# Ollama Model Pilot\n\n> Version: **v1.4.1**  \n> Package / slug: **modelpilot**  \n> Former names: **Model Pilot**, **Ollama Lifecycle Manager**  \n> Positioning: A safety-first operational guide for managing local Ollama model workflows.\n\n---\n\n## v1.4.1 Release Highlights / v1.4.1 发布重点\n\nThis v1.4.1 release is a small activation-focused update based on **v1.4**.\n\nMain changes:\n\n- Added **Activation Rules** so agents know when to use this skill.\n- Made **Real-Task Benchmark** the default workflow when users ask to test, benchmark, compare, evaluate, or select Ollama models, unless they explicitly ask for a quick smoke test only.\n- Added a required benchmark response format to make model-testing results more consistent.\n\n本 v1.4.1 是基于 **v1.4** 的小幅触发规则增强版。\n\n本次更新：\n\n- 新增 **触发规则**，明确哪些场景应调用本 skill。\n- 当用户要求测试、Benchmark、比较、评估或选择 Ollama 模型时，默认使用 **真实任务基准测试**，除非用户明确只要求快速连通性测试。\n- 新增模型测试标准输出格式，让测试结果更稳定、更可复盘。\n\n---\n\n## Author Note / 作者说明\n\nThis skill comes from a real-world local Ollama power-user workflow.\n\nI maintain multiple local models for RAG, document analysis, vision tasks, coding assistance, and role-based AI workflows. As the model pool grows, the hard part is no longer just pulling models. The real challenge is knowing why each model exists, whether it is still referenced by scripts, whether it can be safely removed, whether benchmark results are meaningful, and whether no-think settings actually work.\n\n**Ollama Model Pilot** is designed to be a safety-first operational guide rather than an aggressive automation tool. Its goal is to help users manage local models with more structure and less risk: inspect before changing, keep an inventory, scan references before cleanup, and validate models with real tasks before replacing production workflows.\n\n这个 skill 来自一个真实的本地 Ollama 重度使用场景。\n\n我在本地同时维护多个模型，用于 RAG、文档分析、视觉识别、代码辅助和多角色工作流。模型越来越多以后，真正麻烦的不是下载模型，而是：不知道每个模型为什么存在、是否还被脚本引用、能不能安全删除、Benchmark 结果是否可信、no-think 配置到底有没有生效。\n\n**Ollama Model Pilot** 的重点不是鼓励用户自动执行更多命令，而是帮助用户建立一套更安全、更清晰的本地模型管理流程：先检查，再记录；先扫描引用，再判断删除；先用真实任务测试，再替换生产模型。\n\n---\n\n## Naming Note / 命名说明\n\nThis skill was previously published as **Model Pilot** and earlier as **Ollama Lifecycle Manager**. It is now renamed to **Ollama Model Pilot** to make the Ollama use case explicit while keeping the product name **Model Pilot**.\n\nThe package / slug may remain `modelpilot` for compatibility. The old **Ollama Lifecycle Manager** listing has been merged and redirects to `modelpilot`.\n\n本 skill 早期曾以 **Ollama Lifecycle Manager** 和 **Model Pilot** 名称发布。现在统一命名为 **Ollama Model Pilot**，既保留 Ollama 搜索入口，也保留 Model Pilot 的产品名称。\n\n为保持兼容，安装包名 / slug 可以继续使用 `modelpilot`。旧版 **Ollama Lifecycle Manager** listing 已合并并重定向至 `modelpilot`。\n\n---\n\n## Purpose / 用途\n\n**Ollama Model Pilot** is an operational playbook and checklist for local Ollama model lifecycle management. It helps users safely manage:\n\n- model pulling and import decisions\n- model inventory and role tracking\n- reproducible benchmarking\n- real-task model validation\n- aliases and naming conventions\n- no-think / thinking-control usage\n- script reference scanning\n- safe cleanup of old models\n- ModelScope / mirror fallback and related risks\n- common troubleshooting steps\n\n**Ollama Model Pilot** 是 Ollama 本地模型生命周期管理的**操作指南与检查清单**，用于帮助用户安全地完成：\n\n- 拉取和导入模型前的判断\n- 模型台账和角色记录\n- 可复盘 Benchmark\n- 真实任务基准测试\n- 模型别名和命名规范\n- no-think / thinking 控制\n- 脚本模型引用扫描\n- 旧模型安全清理\n- ModelScope / 镜像回退及风险判断\n- 常见故障排查\n\nTrigger words / 触发词：Ollama, local model, model lifecycle, model inventory, benchmark, model cleanup, model alias, no-think, think false, ModelScope, GGUF, RAG, vision model, coding model, script reference scan, 模型生命周期, 模型台账, 模型清理, 模型测速, 本地模型管理。\n\n---\n\n## Activation Rules / 触发规则\n\nUse this skill when the user asks about any of these Ollama model-management tasks:\n\n- test, benchmark, compare, evaluate, or select a local Ollama model\n- decide whether a model is good enough for a workflow\n- check whether a model can be safely deleted\n- create or review an Ollama model inventory\n- scan project files for model references\n- configure or evaluate no-think / thinking behavior\n- choose a model for RAG, document analysis, coding, vision, or role-based workflows\n\nDefault rule:\n\n> When the user asks to test, benchmark, compare, evaluate, or select an Ollama model, use the **Real-Task Benchmark** workflow unless the user explicitly asks for a quick smoke test only.\n\nA quick smoke test only checks whether the model can run. It is not enough to recommend a model for production use.\n\n当用户提出以下需求时，应使用本 skill：\n\n- 测试、Benchmark、比较、评估或选择本地 Ollama 模型\n- 判断模型是否适合进入正式工作流\n- 判断模型是否可以安全删除\n- 建立或检查 Ollama 模型台账\n- 扫描项目文件中的模型引用\n- 配置或评估 no-think / thinking 行为\n- 为 RAG、文档分析、代码、视觉或多角色工作流选择模型\n\n默认规则：\n\n> 当用户要求测试、Benchmark、比较、评估或选择 Ollama 模型时，除非用户明确只要求快速连通性测试，否则使用 **Real-Task Benchmark / 真实任务基准测试** 流程。\n\n快速连通性测试只能说明模型能运行，不能作为进入正式工作流的依据。\n\n## 0. Safety Boundary / 安全边界\n\nDefault principle:\n\n> **Read first. Explain before changing. Confirm before destructive operations.**\n\n默认原则：\n\n> **只读优先，修改前说明，破坏性操作必须单独确认。**\n\n### 0.1 Read-only operations allowed by default / 默认允许的只读操作\n\nThese commands are generally safe to suggest or run when the user asks for model inspection:\n\n```bash\nollama list\nollama --version\ncurl http://localhost:11434/api/version\ncurl http://localhost:11434/api/tags\npwd\nls\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\ndu -sh ~/.ollama/models/\n```\n\n### 0.2 Operations requiring explicit user confirmation / 必须用户确认后才能执行的操作\n\nThe following operations download, create, copy, modify, upgrade, or delete resources. Explain the impact and wait for user confirmation:\n\n```bash\nollama pull <model>\nollama cp <source_model> <alias>\nollama create <new_model> -f <Modelfile>\nollama rm <model>\nbrew upgrade ollama\n```\n\n### 0.3 Operations that should not be automated / 不应自动执行的操作\n\nDo not automatically perform these operations unless the user clearly asks and confirms:\n\n- Delete a model with `ollama rm`\n- Overwrite an existing Modelfile\n- Modify `~/.zshrc`, `~/.bashrc`, `~/.profile`, or other shell configuration files\n- Modify system-level environment variables\n- Change `OLLAMA_CONTEXT_LENGTH`, `OLLAMA_NUM_PARALLEL`, `OLLAMA_MAX_LOADED_MODELS`\n- Bulk-download multiple large models\n- Replace production workflow model names in scripts\n\n### 0.4 Required confirmation before deleting a model / 删除模型前必须确认\n\nBefore deleting any model, output:\n\n1. Current `ollama list` summary\n2. Model name to delete\n3. Known or inferred usage\n4. Whether script references were found\n5. Whether it is an alias or source model\n6. Possible replacement model\n7. Recovery path after deletion\n8. Explicit confirmation prompt\n\nRecommended confirmation text:\n\n```text\nPlease confirm whether to delete model: <model_name>.\nAfter deletion, recovery may require re-downloading or re-creating the model.\n\n请确认是否删除模型：<model_name>。\n删除后如需恢复，可能需要重新下载或重新 create。\n```\n\n---\n\n## 1. When to Use / 适用场景\n\n### 1.1 Suitable scenarios / 适用场景\n\nUse this skill when the user wants to:\n\n- Manage many local Ollama models\n- Know which model exists for which task\n- Compare local models on the same machine\n- Create stable model aliases\n- Decide whether a model can be safely removed\n- Avoid breaking scripts that reference specific models\n- Handle slow or failed model downloads\n- Use ModelScope or another mirror with caution\n- Create no-think or low-think usage patterns\n- Validate models before using them in RAG, document analysis, coding, vision, or production workflows\n\n### 1.2 Unsuitable scenarios / 不适用场景\n\nDo not use this skill to:\n\n- Automatically decide which models to delete\n- Claim a model is the “best” without testing\n- Auto-modify production scripts\n- Benchmark private or sensitive content without user approval\n- Replace domain-specific professional evaluation\n- Encourage unsafe third-party model sources without checking metadata and trust signals\n\n---\n\n## 2. Environment Check / 环境检查\n\nBefore changing anything, inspect the environment:\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n```\n\nRecord:\n\n```markdown\n## Ollama Environment Record\n\n- Date:\n- Operating system:\n- CPU / GPU / chip:\n- Memory:\n- Ollama CLI version:\n- Ollama server version:\n- Model directory size:\n- Main use cases: chat / RAG / document analysis / coding / vision / other\n- Notes:\n```\n\nIf CLI and server versions appear inconsistent, restart Ollama before testing new models or diagnosing model failures.\n\n---\n\n## 3. Model Inventory / 模型台账\n\nLifecycle management is not “delete models until only a few remain.” The real goal is to know why each model exists.\n\nRecommended file:\n\n```text\nOllama模型台账.md\n```\n\nTemplate:\n\n```markdown\n# Ollama Model Inventory / Ollama 模型台账\n\n| Model | Type | Current use | Script / app reference | Status | Deletable? | Replacement | Last tested | Notes |\n|---|---|---|---|---|---|---|---|---|\n| gemma4:26b | text | long-document analysis | batch_xxx.py | core | no | - | YYYY-MM-DD | stable |\n| qwen3-vl | vision | image understanding | photo_app | task model | no | - | YYYY-MM-DD | keep |\n| gpt-oss:20b | text | experiment | none | testing | maybe | xxx | YYYY-MM-DD | slow |\n```\n\n### 3.1 Model status categories / 模型状态分层\n\nClassify every model into one of these groups:\n\n1. **Core model / 核心模型**  \n   Used by a default workflow. Do not delete.\n\n2. **Task model / 任务模型**  \n   Special use: embedding, vision, coding, long-context, RAG, OCR-related tasks, etc.\n\n3. **Backup model / 备用模型**  \n   Useful when the core model fails, is too slow, or produces weak output.\n\n4. **Testing model / 测试模型**  \n   New model under observation. Revisit after 7–30 days.\n\n5. **Deprecated model / 废弃模型**  \n   No clear use, no script references, has a replacement, safe to remove after confirmation.\n\n---\n\n## 4. Script Reference Scan / 脚本引用扫描\n\nBefore cleaning models, scan project files for hardcoded model references.\n\nFrom the project root:\n\n```bash\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n```\n\nFallback if `rg` is unavailable:\n\n```bash\ngrep -R \"ollama run\\|MODEL_NAME\\|model_name\\|\\\"model\\\"\" . 2>/dev/null\n```\n\nRecord results:\n\n```markdown\n## Model Reference Scan Result / 模型引用扫描结果\n\n| Model | File | Location | Use | Deletion impact |\n|---|---|---|---|---|\n| gemma-doc-nothink | batch_qwen_cards.py | MODEL_NAME | extraction | high |\n| mistral-qa | qa_ask.py | model_name | deep QA | high |\n```\n\nBefore deleting a model, answer:\n\n1. Is it hardcoded in scripts?\n2. Is it an alias or source model?\n3. Is it used for embedding, RAG, vision, coding, or long-context tasks?\n4. Is there a quality- and speed-tested replacement?\n5. Can it be restored easily?\n\n---\n\n## 5. Pulling or Importing a New Model / 拉取或导入新模型\n\n### 5.1 Prefer the official Ollama registry first / 优先使用官方 registry\n\n```bash\nollama pull <model_name>:<tag>\n```\n\nBefore pulling, check:\n\n- model source\n- model size\n- quantization level\n- estimated disk usage\n- hardware memory fit\n- overlap with existing models\n- intended use case\n- test task to run after download\n\n### 5.2 Fallback when download fails / 下载失败后的回退流程\n\nIf the official registry is slow or stuck:\n\n1. Confirm it is not a temporary network issue.\n2. Stop the stuck pull if needed.\n3. Search ModelScope or Hugging Face for the target model.\n4. Verify publisher, filename, quantization, update time, and license.\n5. Pull with a full ModelScope path or manually import GGUF.\n6. Test with the fixed benchmark set before using it in production.\n\nModelScope example:\n\n```bash\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n```\n\n### 5.3 Mirror-source risk warning / 镜像源风险提示\n\nWhen using ModelScope, Hugging Face mirrors, or other third-party sources:\n\n- Use the model page as the source of truth for path and tag.\n- Same model names may be uploaded by unofficial publishers.\n- Template, chat format, license, and quantization may differ from official versions.\n- Pull success does not mean workflow compatibility.\n- Always run fixed benchmark and real-task validation.\n\n---\n\n## 6. Benchmark Standard Process / Benchmark 标准流程\n\nDo not rely only on `ollama run + subprocess` wall-clock timing. Wall time mixes model loading, CLI overhead, output rendering, and generation time.\n\nPrefer the Ollama API and record:\n\n- `total_duration`\n- `load_duration`\n- `prompt_eval_count`\n- `prompt_eval_duration`\n- `eval_count`\n- `eval_duration`\n\nCore speed metric:\n\n```text\ntokens_per_second = eval_count / (eval_duration / 1e9)\n```\n\n### 6.1 Benchmark scenarios / 测试场景\n\nUse at least four scenarios:\n\n1. **Short QA / 短问答**: 100–300 characters output, daily response speed.\n2. **Long writing / 长文写作**: 800–1500 characters output, sustained generation.\n3. **Long summary / 长文摘要**: 5,000–15,000 characters input, context processing.\n4. **Structured output / 结构化输出**: JSON or fixed Markdown, workflow stability.\n\n### 6.2 Cold start and warm start / 冷启动与热启动\n\n- **Cold start** includes model loading time. It reflects first-response experience.\n- **Warm start** runs a warmup prompt before testing. It reflects continuous-use speed.\n\n### 6.3 API benchmark script / API Benchmark 脚本\n\nSave as:\n\n```text\nollama_benchmark.py\n```\n\n```python\n#!/usr/bin/env python3\nimport argparse\nimport json\nimport statistics\nimport time\nimport urllib.request\nfrom typing import Any, Dict, List\n\nOLLAMA_URL = \"http://localhost:11434/api/generate\"\n\n\ndef call_ollama(model: str, prompt: str, think: str | None = None, timeout: int = 600) -> Dict[str, Any]:\n    payload: Dict[str, Any] = {\n        \"model\": model,\n        \"prompt\": prompt,\n        \"stream\": False,\n    }\n\n    if think is not None:\n        if think.lower() == \"false\":\n            payload[\"think\"] = False\n        elif think.lower() == \"true\":\n            payload[\"think\"] = True\n        else:\n            payload[\"think\"] = think\n\n    req = urllib.request.Request(\n        OLLAMA_URL,\n        data=json.dumps(payload).encode(\"utf-8\"),\n        headers={\"Content-Type\": \"application/json\"},\n        method=\"POST\",\n    )\n\n    start = time.time()\n    with urllib.request.urlopen(req, timeout=timeout) as resp:\n        data = json.loads(resp.read().decode(\"utf-8\"))\n    wall_sec = time.time() - start\n\n    eval_count = data.get(\"eval_count\") or 0\n    eval_duration = data.get(\"eval_duration\") or 0\n    prompt_eval_count = data.get(\"prompt_eval_count\") or 0\n    prompt_eval_duration = data.get(\"prompt_eval_duration\") or 0\n    load_duration = data.get(\"load_duration\") or 0\n    total_duration = data.get(\"total_duration\") or 0\n    response = data.get(\"response\", \"\")\n\n    tps = None\n    if eval_count and eval_duration:\n        tps = eval_count / (eval_duration / 1e9)\n\n    prompt_tps = None\n    if prompt_eval_count and prompt_eval_duration:\n        prompt_tps = prompt_eval_count / (prompt_eval_duration / 1e9)\n\n    return {\n        \"model\": model,\n        \"wall_sec\": round(wall_sec, 3),\n        \"total_sec_api\": round(total_duration / 1e9, 3) if total_duration else None,\n        \"load_sec\": round(load_duration / 1e9, 3) if load_duration else None,\n        \"prompt_eval_count\": prompt_eval_count,\n        \"prompt_eval_sec\": round(prompt_eval_duration / 1e9, 3) if prompt_eval_duration else None,\n        \"prompt_tokens_per_sec\": round(prompt_tps, 2) if prompt_tps else None,\n        \"eval_count\": eval_count,\n        \"eval_sec\": round(eval_duration / 1e9, 3) if eval_duration else None,\n        \"tokens_per_sec\": round(tps, 2) if tps else None,\n        \"response_chars\": len(response),\n        \"preview\": response[:160].replace(\"\\n\", \" \") + (\"...\" if len(response) > 160 else \"\"),\n    }\n\n\ndef run_benchmark(models: List[str], prompt: str, rounds: int, warmup: bool, think: str | None) -> None:\n    if warmup:\n        for model in models:\n            try:\n                call_ollama(model, \"请用一句话回答：测试。\", think=think)\n            except Exception as e:\n                print(f\"[WARMUP FAIL] {model}: {e}\")\n\n    for model in models:\n        results = []\n        for i in range(rounds):\n            try:\n                result = call_ollama(model, prompt, think=think)\n                result[\"round\"] = i + 1\n                results.append(result)\n                print(json.dumps(result, ensure_ascii=False))\n            except Exception as e:\n                print(json.dumps({\"model\": model, \"round\": i + 1, \"error\": str(e)}, ensure_ascii=False))\n\n        speeds = [r[\"tokens_per_sec\"] for r in results if r.get(\"tokens_per_sec\") is not None]\n        if speeds:\n            summary = {\n                \"model\": model,\n                \"rounds\": len(speeds),\n                \"tokens_per_sec_avg\": round(statistics.mean(speeds), 2),\n                \"tokens_per_sec_min\": round(min(speeds), 2),\n                \"tokens_per_sec_max\": round(max(speeds), 2),\n            }\n            print(\"[SUMMARY] \" + json.dumps(summary, ensure_ascii=False))\n\n\nif __name__ == \"__main__\":\n    parser = argparse.ArgumentParser(description=\"Ollama API benchmark\")\n    parser.add_argument(\"models\", nargs=\"+\", help=\"Models to benchmark\")\n    parser.add_argument(\"--prompt\", default=\"请写一段约500字的中文说明，介绍本地大模型的实际用途。\")\n    parser.add_argument(\"--rounds\", type=int, default=1)\n    parser.add_argument(\"--warmup\", action=\"store_true\")\n    parser.add_argument(\"--think\", default=None, help=\"true / false / low / medium / high, if supported by model and Ollama\")\n    args = parser.parse_args()\n\n    run_benchmark(args.models, args.prompt, args.rounds, args.warmup, args.think)\n```\n\nUsage:\n\n```bash\npython3 ollama_benchmark.py gemma4:26b qq36 --rounds 2 --warmup\npython3 ollama_benchmark.py qq36-nothink qq36 --think false --rounds 2 --warmup\n```\n\n### 6.4 Quality evaluation is required / 速度测试不能代替质量测试\n\nBefore using a model in a production workflow, test with real or sanitized materials:\n\n- Does it miss important facts?\n- Does it hallucinate?\n- Does it follow fixed output formats?\n- Can it handle long context?\n- Does it over-infer?\n- Does it leak visible reasoning when it should not?\n- Does it match the intended role?\n\n### 6.5 Real-Task Benchmark / 真实任务基准测试\n\nDo not test models only with one generic prompt. Ollama users usually care less about abstract speed and more about this question:\n\n> **Does this model perform well in my real task without breaking format, facts, or workflow stability?**\n\n建议为常用模型建立一组小型、固定、可复用的测试集。每次测试使用同一批任务、同一套评分维度、同一份结果记录模板。\n\nRecommended task set:\n\n1. **Short QA / 短问答**  \n   Tests basic understanding and daily response speed.\n\n2. **Long Summary / 长文摘要**  \n   Input 5,000–15,000 characters. Tests context handling, fact retention, and compression.\n\n3. **Structured Output / 结构化输出**  \n   Requires JSON, Markdown table, or fixed fields. Tests workflow reliability.\n\n4. **Role Task / 真实角色任务**  \n   Uses the user’s own real workflow task or a sanitized sample. Tests production fit.\n\nResult template:\n\n```markdown\n## Real-Task Benchmark Result\n\n- Date:\n- Tester:\n- Device:\n- Ollama version:\n- Model:\n- Quantization:\n- Think setting: true / false / low / medium / high / unset\n- Task type: Short QA / Long Summary / Structured Output / Role Task\n- prompt_eval_count:\n- eval_count:\n- load_duration:\n- prompt_eval_duration:\n- eval_duration:\n- tokens_per_second:\n- total_duration:\n- format_pass: yes / no\n- quality_score: 1–5\n- hallucination_risk: low / medium / high\n- production_ready: yes / no / observe\n- Notes:\n```\n\nQuality score:\n\n| Score | Meaning | Standard |\n|---:|---|---|\n| 5 | Production-ready | Acceptable speed, stable facts, stable format, not weaker than the current model |\n| 4 | Small-scope trial | Good quality but needs more observation |\n| 3 | Backup / testing only | Useful but has clear weaknesses |\n| 2 | Not recommended | Serious problems in speed, facts, or format |\n| 1 | Reject | Fails the task or frequently breaks output |\n\nKey rule:\n\n> A model should not be promoted into a production workflow only because it is fast. It should pass at least one real-task benchmark that matches the user’s actual use case.\n\n中文原则：\n\n> 不能因为一个模型跑得快，就把它放进正式工作流。至少要通过一个贴近真实使用场景的任务测试，才能替换生产模型。\n\nDecision guide:\n\n```text\nFast but unstable: keep as test only.\nHigh quality but slow: use as deep-analysis or backup model.\nFormat unstable: do not use for RAG, automation, or structured output.\nWins real-task benchmark: replace gradually and observe for 7–30 days.\n```\n\n### 6.6 Required Benchmark Response Format / 模型测试标准输出格式\n\nWhen using Real-Task Benchmark, respond with this structure:\n\n1. Benchmark goal\n2. Model and environment\n3. Test tasks selected\n4. Speed metrics\n5. Output quality\n6. Format stability\n7. Workflow fit\n8. Risks or limitations\n9. Promotion decision\n10. Recommended next action\n\nClearly distinguish:\n\n- smoke test\n- speed benchmark\n- real-task benchmark\n- production promotion decision\n\n使用真实任务基准测试时，按以下结构输出：\n\n1. 测试目标\n2. 模型与环境\n3. 选用的测试任务\n4. 速度指标\n5. 输出质量\n6. 格式稳定性\n7. 工作流适配度\n8. 风险或限制\n9. 是否建议进入正式工作流\n10. 下一步建议\n\n必须区分：快速连通性测试、速度测试、真实任务基准测试、是否进入正式工作流的决策。\n\n---\n\n## 7. Aliases and Naming / 别名与命名\n\n### 7.1 Ollama model aliases / Ollama 模型别名\n\n```bash\nollama cp <source_model> <alias_name>\n```\n\nExample:\n\n```bash\nollama cp qwen3.6:35b-a3b-q4_k_m qq36\n```\n\nNotes:\n\n- `ollama cp` usually creates an alias sharing underlying model storage.\n- Deleting an alias is not the same as deleting the source model.\n- Before deleting a source model, confirm whether aliases depend on it.\n- Alias names should be short, stable, and meaningful.\n\n### 7.2 Shell aliases / Shell 命令别名\n\nOnly suggest commands by default. Do not automatically write to shell config.\n\n```bash\nalias qq36='ollama run qq36'\nalias g26='ollama run gemma4:26b'\n```\n\nIf the user wants to persist them:\n\n```bash\necho \"alias qq36='ollama run qq36'\" >> ~/.zshrc\nsource ~/.zshrc\n```\n\nWarn the user before modifying shell configuration files.\n\n---\n\n## 8. No-Think / Thinking Control / 思考模式控制规范\n\n### 8.1 Core principle / 核心原则\n\nA SYSTEM prompt can ask a model not to show reasoning, but it does not necessarily disable internal thinking.\n\nSYSTEM 提示词只能要求模型“不输出思考过程”，不等于底层关闭 thinking。\n\nIf Ollama and the model support the `think` parameter, prefer explicit API control:\n\n```bash\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"请总结这段材料。\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\nIf thinking levels are supported, test:\n\n```json\n\"think\": \"low\"\n```\n\nor:\n\n```json\n\"think\": \"medium\"\n```\n\nAlways verify through actual output and API behavior.\n\n### 8.2 Modelfile is only a weak constraint / Modelfile 只是弱约束\n\nFor models that do not support `think`, a Modelfile can reduce visible reasoning but cannot guarantee true no-think behavior.\n\n```text\nFROM <model_name>:<tag>\nSYSTEM \"直接输出最终答案。不要输出思考过程、推理草稿或隐藏分析。\"\nPARAMETER temperature 0.4\n```\n\nCreate:\n\n```bash\nollama create <model_name>-nothink -f Modelfile.<model_name>-nothink\n```\n\nRecommended naming:\n\n```text\nModelfile.<source_model>.<purpose>.nothink\n```\n\nDo not overwrite existing Modelfiles.\n\n---\n\n## 9. Cleanup Workflow / 旧模型清理\n\n### 9.1 Inspect models and disk usage / 查看模型与占用\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n### 9.2 Cleanup criteria / 清理判断标准\n\nDo not mechanically keep only 3–5 models. A model is a cleanup candidate only when most of these are true:\n\n- No clear use\n- No script references\n- Not an embedding / vision / coding / RAG / long-context model\n- Not used in the last 30 days, or confirmed no longer needed\n- Benchmark is clearly worse than alternatives\n- A tested replacement exists\n- It can be re-downloaded or rebuilt\n\n### 9.3 Pre-delete checklist / 删除前检查模板\n\n```markdown\n## Pre-delete Model Check / 删除模型前检查\n\n- Model to delete:\n- Current use:\n- In inventory: yes / no\n- Script reference found: yes / no\n- Is alias: yes / no / unknown\n- Replacement model:\n- Deletion risk: low / medium / high\n- Recovery path: re-pull / re-create / unavailable\n- Recommendation: delete / hold / keep\n```\n\nDelete only after confirmation:\n\n```bash\nollama rm <model_name>\n```\n\nVerify after deletion:\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n---\n\n## 10. Quantization Selection / 量化级别选择指南\n\n| Quantization | Size | Quality | Speed / memory | Use case |\n|---|---:|---|---|---|\n| Q3_K_M | small | basic | low pressure | memory-constrained, rough tasks |\n| Q4_K_M | medium | good | balanced | default daily choice |\n| Q5_K_M | larger | better | more pressure | quality-sensitive tasks |\n| Q8_0 | large | close to original | high pressure | quality-first and enough memory |\n\nRules:\n\n- Chat, summary, light workflows: start with Q4_K_M.\n- Higher-quality long-form or complex analysis: test Q5_K_M or Q8_0.\n- Low-memory devices: try smaller models or Q3_K_M.\n- Compare quantizations with the same benchmark set.\n\n---\n\n## 11. Troubleshooting / 常见故障\n\n### 11.1 Ollama version too old / Ollama 版本过旧\n\nSymptoms:\n\n```text\nunable to load model\n```\n\nPossible causes:\n\n- Ollama does not support the model format\n- CLI and server versions differ\n- Ollama server has not restarted\n\nCheck:\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\n```\n\nRestart or upgrade based on the installation method.\n\n### 11.2 Model download stuck / 模型下载卡住\n\nOrder of handling:\n\n1. Check network.\n2. Retry `ollama pull`; Ollama may resume.\n3. If official source remains stuck, consider ModelScope or Hugging Face.\n4. After using another source, validate model quality and format.\n\n### 11.3 Slow loading or memory pressure / 加载慢或内存不足\n\nCheck:\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\nSuggested actions:\n\n- Use a smaller model or lower quantization.\n- Reduce concurrency.\n- Avoid loading multiple large models at once.\n- Do not blindly raise `OLLAMA_CONTEXT_LENGTH`.\n- Use long context only when the task requires it.\n\n### 11.4 Output format unstable / 输出格式不稳定\n\nPossible causes:\n\n- Model weak at structured output\n- Temperature too high\n- Prompt too loose\n- Context too long and instructions diluted\n\nActions:\n\n- Lower temperature.\n- Use a stricter template.\n- Test with short structured samples.\n- Switch model if needed.\n\n---\n\n## 12. Recommended Workflows / 推荐工作流\n\n### 12.1 New model adoption / 新模型进入流程\n\n```text\nDiscover model → Check source → Define use case → Pull/import → Record environment → Benchmark → Real-task test → Add to inventory → Decide whether to bind to scripts\n```\n\n### 12.2 Model replacement / 模型替换流程\n\n```text\nTest candidate → Compare with current model → Check quality and format → Replace in small scope → Keep old model during observation → Update inventory → Decide whether to delete old model\n```\n\n### 12.3 Cleanup / 模型清理流程\n\n```text\nollama list → Inventory check → Script reference scan → Replacement check → Deletion recommendation → User confirmation → Delete → Verify disk usage → Update inventory\n```\n\n---\n\n## 13. Minimal Command Set / 最小可复制命令清单\n\n```bash\n# Environment\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n\n# Pull model\nollama pull <model_name>:<tag>\n\n# ModelScope fallback\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n\n# Create Ollama alias\nollama cp <source_model> <alias_name>\n\n# Scan script references\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n\n# Delete model, confirmation required\nollama rm <model_name>\n\n# API no-think test\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"请用一句话回答：测试。\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\n---\n\n## 14. Response Requirements / 输出要求\n\nWhen the user asks to manage Ollama models, respond with:\n\n1. Current judgment\n2. Safety risks\n3. Read-only checks to run first\n4. Whether confirmation is needed\n5. Next command or next decision\n\nWhen the user asks to clean up models, respond with:\n\n1. Model grouping\n2. Models that must not be deleted\n3. Models to observe\n4. Deletion candidates\n5. Script-reference scan command\n6. Explicit deletion confirmation prompt\n\nWhen the user asks to test models, respond with:\n\n1. Test goal\n2. Test scenarios\n3. Benchmark command\n4. Quality dimensions\n5. Production-readiness judgment\n\n---\n\n## 15. Version Notes / 版本说明\n\n### v1.4.1\n\n- Added Activation Rules so agents know when to use this skill.\n- Made Real-Task Benchmark the default workflow for model testing, comparison, evaluation, and selection unless the user asks for smoke testing only.\n- Added Required Benchmark Response Format for consistent results.\n\n### v1.4\n\n- Upgraded from the previous **v1.3 Model Pilot** release.\n- Renamed the maintained skill to **Ollama Model Pilot** while keeping `modelpilot` as the recommended package / slug for compatibility.\n- Added a clear v1.4 release highlight section near the top of the skill file.\n- Promoted **Real-Task Benchmark / 真实任务基准测试** into a core model-testing workflow, not just a supplementary note.\n- Added the production replacement rule: a model should not replace a workflow only because it is fast; it must pass at least one real-task benchmark matching the user’s actual use case.\n- Clarified legacy names: **Ollama Lifecycle Manager** and **Model Pilot**.\n- Clarified that old listings should be treated as legacy migration entries pointing to this maintained version.\n- Preserved the safety-first workflow: inventory first, script-reference scan before cleanup, confirmation before deletion, and cautious no-think handling.\n\n### v1.3\n\n- Released under the **Model Pilot** naming line.\n- Added / retained author note and safety-first positioning.\n- Included model inventory, script reference scanning, API benchmark, no-think rules, cleanup confirmation, and mirror-source risk warnings.\n\n---\n\n## 16. License / 许可说明\n\nIf this skill is derived from another published skill, confirm the original license and attribution requirements before distributing.\n\n如果本 skill 基于其他已发布 skill 改写或增强，请在发布前确认原始许可、署名和再分发要求。\n\nFile v1.4.1:README.md\n\n# Ollama Model Pilot v1.4.1\n\nA safety-first operational guide for managing local Ollama model workflows.\n\nThis v1.4.1 release is a small activation-focused update based on v1.4.\n\n## What it does\n\nOllama Model Pilot helps local Ollama users manage:\n\n- model inventory and role tracking\n- safe cleanup with deletion confirmation\n- script reference scanning before removing models\n- API-based benchmark\n- Real-Task Benchmark for production model validation\n- no-think / thinking-control usage\n- aliases and naming conventions\n- ModelScope / mirror fallback risk checks\n- common troubleshooting\n\n## Main v1.4.1 update\n\nv1.4.1 adds clearer **Activation Rules**.\n\nWhen users ask to test, benchmark, compare, evaluate, or select an Ollama model, agents should use the **Real-Task Benchmark** workflow by default, unless the user explicitly asks for a quick smoke test only.\n\nIt also adds a required benchmark response format so results are easier to compare and review.\n\n## Install from ClawHub\n\n```bash\nopenclaw skills install modelpilot\n```\n\n## Naming\n\nThe maintained name is **Ollama Model Pilot**.\n\nEarlier names: **Model Pilot**, **Ollama Lifecycle Manager**.\n\nThe old **Ollama Lifecycle Manager** listing has been merged and redirects to `modelpilot`.\n\n## Version\n\nCurrent version: **v1.4.1**\n\nPrevious version: **v1.4**\n\n## Files\n\n- `SKILL.md` — main skill file\n- `README.md` — publishing note\n- `CHANGELOG.md` — version history\n\nFile v1.4.1:_meta.json\n\n{\n  \"ownerId\": \"kn7fj0qnbect2fmp59jj0vx34587hvns\",\n  \"slug\": \"modelpilot\",\n  \"version\": \"1.4.1\",\n  \"publishedAt\": 1780037005998\n}\n\nFile v1.4.1:CHANGELOG.md\n\n# Changelog\n\n## v1.4.1\n\n- Added **Activation Rules** so agents know when to use this skill.\n- Made **Real-Task Benchmark** the default workflow when users ask to test, benchmark, compare, evaluate, or select Ollama models, unless they explicitly ask for a quick smoke test only.\n- Added **Required Benchmark Response Format** for more consistent model-testing results.\n- Kept the maintained name as **Ollama Model Pilot** and package / slug as `modelpilot`.\n- Clarified that the old **Ollama Lifecycle Manager** listing has been merged and redirects to `modelpilot`.\n\n## v1.4\n\n- Upgraded from the previous **v1.3 Model Pilot** release.\n- Renamed the maintained skill to **Ollama Model Pilot**.\n- Kept `modelpilot` as the recommended package / slug for compatibility.\n- Promoted **Real-Task Benchmark / 真实任务基准测试** into the core model-testing workflow.\n- Added the production replacement rule: do not replace a workflow only because a model is faster; require at least one matching real-task benchmark pass.\n- Clarified legacy names: **Ollama Lifecycle Manager** and **Model Pilot**.\n- Preserved safety-first model management: inventory, script-reference scanning, deletion confirmation, no-think caution, and mirror-source risk checks.\n\n## v1.3\n\n- Released under the **Model Pilot** naming line.\n- Included author note and safety-first positioning.\n- Included model inventory, script reference scanning, API benchmark, no-think rules, cleanup confirmation, and mirror-source risk warnings.\n\nFile v1.4.1:skill-card.md\n\n## Description: <br>\nA safety-first operational guide for managing local Ollama model workflows, including inventory, benchmarking, aliases, cleanup checks, and troubleshooting. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[patmenciu](https://clawhub.ai/user/patmenciu) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and local Ollama users use this skill to inspect model environments, maintain model inventories, benchmark and compare models, check script references, and confirm cleanup steps before changing local workflows. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Model pulls, alias creation, shell configuration edits, upgrades, and model deletion can change local resources or workflows. <br>\nMitigation: Review proposed commands before running them and require explicit confirmation before downloads, modifications, upgrades, or deletion. <br>\nRisk: Deleting a model may break scripts, aliases, RAG workflows, document analysis, vision tasks, coding tasks, or role-based workflows that still reference it. <br>\nMitigation: Run the script-reference scan, record known usage, identify whether the model is an alias or source model, and confirm a recovery path before deletion. <br>\nRisk: Third-party or mirror model sources may differ in publisher, template, chat format, license, quantization, or workflow compatibility. <br>\nMitigation: Prefer official sources first; when using mirrors, verify publisher and model metadata and run fixed benchmarks plus real-task validation before production use. <br>\n\n\n## Reference(s): <br>\n- [ClawHub Skill Page](https://clawhub.ai/patmenciu/modelpilot) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Guidance, Markdown, Code, Shell commands, Configuration] <br>\n**Output Format:** [Markdown with inline bash, Python, JSON, and checklist templates] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Includes confirmation prompts for destructive operations and structured benchmark result templates.] <br>\n\n## Skill Version(s): <br>\n1.4.1 (source: server release evidence, README, SKILL.md, CHANGELOG) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.4.0: 5 files, 15819 bytes\n\nFiles: CHANGELOG.md (1096b), README.md (1876b), skill-card.md (2538b), SKILL.md (30924b), _meta.json (129b)\n\nFile v1.4.0:SKILL.md\n\n# Ollama Model Pilot\n\n> Version: **v1.4**  \n> Package / slug: **modelpilot**  \n> Former names: **Model Pilot**, **Ollama Lifecycle Manager**  \n> Positioning: A safety-first operational guide for managing local Ollama model workflows.\n\n---\n\n## v1.4 Release Highlights / v1.4 发布重点\n\nThis v1.4 release is upgraded from the previous **v1.3 Model Pilot** release. The maintained name is now **Ollama Model Pilot**.\n\nMain upgrade in this release:\n\n- **Real-Task Benchmark is promoted to a core model-testing workflow.** Model testing should not stop at tokens-per-second. A model should be validated with fixed, repeatable tasks that match real use cases: Short QA, Long Summary, Structured Output, and Role Task.\n- **Production replacement rule is now explicit.** A model should not replace an existing workflow only because it is faster; it should pass at least one real-task benchmark with acceptable quality, fact stability, and output-format reliability.\n- **Legacy naming is clarified.** Earlier names include **Ollama Lifecycle Manager** and **Model Pilot**. The maintained product name is **Ollama Model Pilot**, while the package / slug can remain `modelpilot` for compatibility.\n\n本 v1.4 版本基于上一版 **v1.3 Model Pilot** 升级，维护名称统一为 **Ollama Model Pilot**。\n\n本次核心升级：\n\n- **把“真实任务基准测试”提升为模型测试核心流程。** 模型测试不能只看 tokens/s，而要用固定、可复用、贴近真实场景的任务集测试：短问答、长文摘要、结构化输出、真实角色任务。\n- **明确生产替换规则。** 不能因为一个模型更快，就直接替换正式工作流；至少要通过一个真实任务基准测试，并在质量、事实稳定性和格式稳定性上达标。\n- **理清历史命名。** 早期名称包括 **Ollama Lifecycle Manager** 和 **Model Pilot**；当前维护名称为 **Ollama Model Pilot**，安装包名 / slug 可继续保留 `modelpilot`。\n\n---\n\n## Author Note / 作者说明\n\nThis skill comes from a real-world local Ollama power-user workflow.\n\nI maintain multiple local models for RAG, document analysis, vision tasks, coding assistance, and role-based AI workflows. As the model pool grows, the hard part is no longer just pulling models. The real challenge is knowing why each model exists, whether it is still referenced by scripts, whether it can be safely removed, whether benchmark results are meaningful, and whether no-think settings actually work.\n\n**Ollama Model Pilot** is designed to be a safety-first operational guide rather than an aggressive automation tool. Its goal is to help users manage local models with more structure and less risk: inspect before changing, keep an inventory, scan references before cleanup, and validate models with real tasks before replacing production workflows.\n\n这个 skill 来自一个真实的本地 Ollama 重度使用场景。\n\n我在本地同时维护多个模型，用于 RAG、文档分析、视觉识别、代码辅助和多角色工作流。模型越来越多以后，真正麻烦的不是下载模型，而是：不知道每个模型为什么存在、是否还被脚本引用、能不能安全删除、Benchmark 结果是否可信、no-think 配置到底有没有生效。\n\n**Ollama Model Pilot** 的重点不是鼓励用户自动执行更多命令，而是帮助用户建立一套更安全、更清晰的本地模型管理流程：先检查，再记录；先扫描引用，再判断删除；先用真实任务测试，再替换生产模型。\n\n---\n\n## Naming Note / 命名说明\n\nThis skill was previously published as **Model Pilot** and earlier as **Ollama Lifecycle Manager**. It is now renamed to **Ollama Model Pilot** to make the Ollama use case explicit while keeping the product name **Model Pilot**.\n\nThe package / slug may remain `modelpilot` for compatibility. If an older listing named **Ollama Lifecycle Manager** still exists, it should be treated as a legacy entry and should point users to this maintained version.\n\n本 skill 早期曾以 **Ollama Lifecycle Manager** 和 **Model Pilot** 名称发布。现在统一命名为 **Ollama Model Pilot**，既保留 Ollama 搜索入口，也保留 Model Pilot 的产品名称。\n\n为保持兼容，安装包名 / slug 可以继续使用 `modelpilot`。如果旧版 **Ollama Lifecycle Manager** 页面仍然存在，应作为历史入口处理，并引导用户使用本维护版本。\n\n---\n\n## Purpose / 用途\n\n**Ollama Model Pilot** is an operational playbook and checklist for local Ollama model lifecycle management. It helps users safely manage:\n\n- model pulling and import decisions\n- model inventory and role tracking\n- reproducible benchmarking\n- real-task model validation\n- aliases and naming conventions\n- no-think / thinking-control usage\n- script reference scanning\n- safe cleanup of old models\n- ModelScope / mirror fallback and related risks\n- common troubleshooting steps\n\n**Ollama Model Pilot** 是 Ollama 本地模型生命周期管理的**操作指南与检查清单**，用于帮助用户安全地完成：\n\n- 拉取和导入模型前的判断\n- 模型台账和角色记录\n- 可复盘 Benchmark\n- 真实任务基准测试\n- 模型别名和命名规范\n- no-think / thinking 控制\n- 脚本模型引用扫描\n- 旧模型安全清理\n- ModelScope / 镜像回退及风险判断\n- 常见故障排查\n\nTrigger words / 触发词：Ollama, local model, model lifecycle, model inventory, benchmark, model cleanup, model alias, no-think, think false, ModelScope, GGUF, RAG, vision model, coding model, script reference scan, 模型生命周期, 模型台账, 模型清理, 模型测速, 本地模型管理。\n\n---\n\n## 0. Safety Boundary / 安全边界\n\nDefault principle:\n\n> **Read first. Explain before changing. Confirm before destructive operations.**\n\n默认原则：\n\n> **只读优先，修改前说明，破坏性操作必须单独确认。**\n\n### 0.1 Read-only operations allowed by default / 默认允许的只读操作\n\nThese commands are generally safe to suggest or run when the user asks for model inspection:\n\n```bash\nollama list\nollama --version\ncurl http://localhost:11434/api/version\ncurl http://localhost:11434/api/tags\npwd\nls\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\ndu -sh ~/.ollama/models/\n```\n\n### 0.2 Operations requiring explicit user confirmation / 必须用户确认后才能执行的操作\n\nThe following operations download, create, copy, modify, upgrade, or delete resources. Explain the impact and wait for user confirmation:\n\n```bash\nollama pull <model>\nollama cp <source_model> <alias>\nollama create <new_model> -f <Modelfile>\nollama rm <model>\nbrew upgrade ollama\n```\n\n### 0.3 Operations that should not be automated / 不应自动执行的操作\n\nDo not automatically perform these operations unless the user clearly asks and confirms:\n\n- Delete a model with `ollama rm`\n- Overwrite an existing Modelfile\n- Modify `~/.zshrc`, `~/.bashrc`, `~/.profile`, or other shell configuration files\n- Modify system-level environment variables\n- Change `OLLAMA_CONTEXT_LENGTH`, `OLLAMA_NUM_PARALLEL`, `OLLAMA_MAX_LOADED_MODELS`\n- Bulk-download multiple large models\n- Replace production workflow model names in scripts\n\n### 0.4 Required confirmation before deleting a model / 删除模型前必须确认\n\nBefore deleting any model, output:\n\n1. Current `ollama list` summary\n2. Model name to delete\n3. Known or inferred usage\n4. Whether script references were found\n5. Whether it is an alias or source model\n6. Possible replacement model\n7. Recovery path after deletion\n8. Explicit confirmation prompt\n\nRecommended confirmation text:\n\n```text\nPlease confirm whether to delete model: <model_name>.\nAfter deletion, recovery may require re-downloading or re-creating the model.\n\n请确认是否删除模型：<model_name>。\n删除后如需恢复，可能需要重新下载或重新 create。\n```\n\n---\n\n## 1. When to Use / 适用场景\n\n### 1.1 Suitable scenarios / 适用场景\n\nUse this skill when the user wants to:\n\n- Manage many local Ollama models\n- Know which model exists for which task\n- Compare local models on the same machine\n- Create stable model aliases\n- Decide whether a model can be safely removed\n- Avoid breaking scripts that reference specific models\n- Handle slow or failed model downloads\n- Use ModelScope or another mirror with caution\n- Create no-think or low-think usage patterns\n- Validate models before using them in RAG, document analysis, coding, vision, or production workflows\n\n### 1.2 Unsuitable scenarios / 不适用场景\n\nDo not use this skill to:\n\n- Automatically decide which models to delete\n- Claim a model is the “best” without testing\n- Auto-modify production scripts\n- Benchmark private or sensitive content without user approval\n- Replace domain-specific professional evaluation\n- Encourage unsafe third-party model sources without checking metadata and trust signals\n\n---\n\n## 2. Environment Check / 环境检查\n\nBefore changing anything, inspect the environment:\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n```\n\nRecord:\n\n```markdown\n## Ollama Environment Record\n\n- Date:\n- Operating system:\n- CPU / GPU / chip:\n- Memory:\n- Ollama CLI version:\n- Ollama server version:\n- Model directory size:\n- Main use cases: chat / RAG / document analysis / coding / vision / other\n- Notes:\n```\n\nIf CLI and server versions appear inconsistent, restart Ollama before testing new models or diagnosing model failures.\n\n---\n\n## 3. Model Inventory / 模型台账\n\nLifecycle management is not “delete models until only a few remain.” The real goal is to know why each model exists.\n\nRecommended file:\n\n```text\nOllama模型台账.md\n```\n\nTemplate:\n\n```markdown\n# Ollama Model Inventory / Ollama 模型台账\n\n| Model | Type | Current use | Script / app reference | Status | Deletable? | Replacement | Last tested | Notes |\n|---|---|---|---|---|---|---|---|---|\n| gemma4:26b | text | long-document analysis | batch_xxx.py | core | no | - | YYYY-MM-DD | stable |\n| qwen3-vl | vision | image understanding | photo_app | task model | no | - | YYYY-MM-DD | keep |\n| gpt-oss:20b | text | experiment | none | testing | maybe | xxx | YYYY-MM-DD | slow |\n```\n\n### 3.1 Model status categories / 模型状态分层\n\nClassify every model into one of these groups:\n\n1. **Core model / 核心模型**  \n   Used by a default workflow. Do not delete.\n\n2. **Task model / 任务模型**  \n   Special use: embedding, vision, coding, long-context, RAG, OCR-related tasks, etc.\n\n3. **Backup model / 备用模型**  \n   Useful when the core model fails, is too slow, or produces weak output.\n\n4. **Testing model / 测试模型**  \n   New model under observation. Revisit after 7–30 days.\n\n5. **Deprecated model / 废弃模型**  \n   No clear use, no script references, has a replacement, safe to remove after confirmation.\n\n---\n\n## 4. Script Reference Scan / 脚本引用扫描\n\nBefore cleaning models, scan project files for hardcoded model references.\n\nFrom the project root:\n\n```bash\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n```\n\nFallback if `rg` is unavailable:\n\n```bash\ngrep -R \"ollama run\\|MODEL_NAME\\|model_name\\|\\\"model\\\"\" . 2>/dev/null\n```\n\nRecord results:\n\n```markdown\n## Model Reference Scan Result / 模型引用扫描结果\n\n| Model | File | Location | Use | Deletion impact |\n|---|---|---|---|---|\n| gemma-doc-nothink | batch_qwen_cards.py | MODEL_NAME | extraction | high |\n| mistral-qa | qa_ask.py | model_name | deep QA | high |\n```\n\nBefore deleting a model, answer:\n\n1. Is it hardcoded in scripts?\n2. Is it an alias or source model?\n3. Is it used for embedding, RAG, vision, coding, or long-context tasks?\n4. Is there a quality- and speed-tested replacement?\n5. Can it be restored easily?\n\n---\n\n## 5. Pulling or Importing a New Model / 拉取或导入新模型\n\n### 5.1 Prefer the official Ollama registry first / 优先使用官方 registry\n\n```bash\nollama pull <model_name>:<tag>\n```\n\nBefore pulling, check:\n\n- model source\n- model size\n- quantization level\n- estimated disk usage\n- hardware memory fit\n- overlap with existing models\n- intended use case\n- test task to run after download\n\n### 5.2 Fallback when download fails / 下载失败后的回退流程\n\nIf the official registry is slow or stuck:\n\n1. Confirm it is not a temporary network issue.\n2. Stop the stuck pull if needed.\n3. Search ModelScope or Hugging Face for the target model.\n4. Verify publisher, filename, quantization, update time, and license.\n5. Pull with a full ModelScope path or manually import GGUF.\n6. Test with the fixed benchmark set before using it in production.\n\nModelScope example:\n\n```bash\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n```\n\n### 5.3 Mirror-source risk warning / 镜像源风险提示\n\nWhen using ModelScope, Hugging Face mirrors, or other third-party sources:\n\n- Use the model page as the source of truth for path and tag.\n- Same model names may be uploaded by unofficial publishers.\n- Template, chat format, license, and quantization may differ from official versions.\n- Pull success does not mean workflow compatibility.\n- Always run fixed benchmark and real-task validation.\n\n---\n\n## 6. Benchmark Standard Process / Benchmark 标准流程\n\nDo not rely only on `ollama run + subprocess` wall-clock timing. Wall time mixes model loading, CLI overhead, output rendering, and generation time.\n\nPrefer the Ollama API and record:\n\n- `total_duration`\n- `load_duration`\n- `prompt_eval_count`\n- `prompt_eval_duration`\n- `eval_count`\n- `eval_duration`\n\nCore speed metric:\n\n```text\ntokens_per_second = eval_count / (eval_duration / 1e9)\n```\n\n### 6.1 Benchmark scenarios / 测试场景\n\nUse at least four scenarios:\n\n1. **Short QA / 短问答**: 100–300 characters output, daily response speed.\n2. **Long writing / 长文写作**: 800–1500 characters output, sustained generation.\n3. **Long summary / 长文摘要**: 5,000–15,000 characters input, context processing.\n4. **Structured output / 结构化输出**: JSON or fixed Markdown, workflow stability.\n\n### 6.2 Cold start and warm start / 冷启动与热启动\n\n- **Cold start** includes model loading time. It reflects first-response experience.\n- **Warm start** runs a warmup prompt before testing. It reflects continuous-use speed.\n\n### 6.3 API benchmark script / API Benchmark 脚本\n\nSave as:\n\n```text\nollama_benchmark.py\n```\n\n```python\n#!/usr/bin/env python3\nimport argparse\nimport json\nimport statistics\nimport time\nimport urllib.request\nfrom typing import Any, Dict, List\n\nOLLAMA_URL = \"http://localhost:11434/api/generate\"\n\n\ndef call_ollama(model: str, prompt: str, think: str | None = None, timeout: int = 600) -> Dict[str, Any]:\n    payload: Dict[str, Any] = {\n        \"model\": model,\n        \"prompt\": prompt,\n        \"stream\": False,\n    }\n\n    if think is not None:\n        if think.lower() == \"false\":\n            payload[\"think\"] = False\n        elif think.lower() == \"true\":\n            payload[\"think\"] = True\n        else:\n            payload[\"think\"] = think\n\n    req = urllib.request.Request(\n        OLLAMA_URL,\n        data=json.dumps(payload).encode(\"utf-8\"),\n        headers={\"Content-Type\": \"application/json\"},\n        method=\"POST\",\n    )\n\n    start = time.time()\n    with urllib.request.urlopen(req, timeout=timeout) as resp:\n        data = json.loads(resp.read().decode(\"utf-8\"))\n    wall_sec = time.time() - start\n\n    eval_count = data.get(\"eval_count\") or 0\n    eval_duration = data.get(\"eval_duration\") or 0\n    prompt_eval_count = data.get(\"prompt_eval_count\") or 0\n    prompt_eval_duration = data.get(\"prompt_eval_duration\") or 0\n    load_duration = data.get(\"load_duration\") or 0\n    total_duration = data.get(\"total_duration\") or 0\n    response = data.get(\"response\", \"\")\n\n    tps = None\n    if eval_count and eval_duration:\n        tps = eval_count / (eval_duration / 1e9)\n\n    prompt_tps = None\n    if prompt_eval_count and prompt_eval_duration:\n        prompt_tps = prompt_eval_count / (prompt_eval_duration / 1e9)\n\n    return {\n        \"model\": model,\n        \"wall_sec\": round(wall_sec, 3),\n        \"total_sec_api\": round(total_duration / 1e9, 3) if total_duration else None,\n        \"load_sec\": round(load_duration / 1e9, 3) if load_duration else None,\n        \"prompt_eval_count\": prompt_eval_count,\n        \"prompt_eval_sec\": round(prompt_eval_duration / 1e9, 3) if prompt_eval_duration else None,\n        \"prompt_tokens_per_sec\": round(prompt_tps, 2) if prompt_tps else None,\n        \"eval_count\": eval_count,\n        \"eval_sec\": round(eval_duration / 1e9, 3) if eval_duration else None,\n        \"tokens_per_sec\": round(tps, 2) if tps else None,\n        \"response_chars\": len(response),\n        \"preview\": response[:160].replace(\"\\n\", \" \") + (\"...\" if len(response) > 160 else \"\"),\n    }\n\n\ndef run_benchmark(models: List[str], prompt: str, rounds: int, warmup: bool, think: str | None) -> None:\n    if warmup:\n        for model in models:\n            try:\n                call_ollama(model, \"请用一句话回答：测试。\", think=think)\n            except Exception as e:\n                print(f\"[WARMUP FAIL] {model}: {e}\")\n\n    for model in models:\n        results = []\n        for i in range(rounds):\n            try:\n                result = call_ollama(model, prompt, think=think)\n                result[\"round\"] = i + 1\n                results.append(result)\n                print(json.dumps(result, ensure_ascii=False))\n            except Exception as e:\n                print(json.dumps({\"model\": model, \"round\": i + 1, \"error\": str(e)}, ensure_ascii=False))\n\n        speeds = [r[\"tokens_per_sec\"] for r in results if r.get(\"tokens_per_sec\") is not None]\n        if speeds:\n            summary = {\n                \"model\": model,\n                \"rounds\": len(speeds),\n                \"tokens_per_sec_avg\": round(statistics.mean(speeds), 2),\n                \"tokens_per_sec_min\": round(min(speeds), 2),\n                \"tokens_per_sec_max\": round(max(speeds), 2),\n            }\n            print(\"[SUMMARY] \" + json.dumps(summary, ensure_ascii=False))\n\n\nif __name__ == \"__main__\":\n    parser = argparse.ArgumentParser(description=\"Ollama API benchmark\")\n    parser.add_argument(\"models\", nargs=\"+\", help=\"Models to benchmark\")\n    parser.add_argument(\"--prompt\", default=\"请写一段约500字的中文说明，介绍本地大模型的实际用途。\")\n    parser.add_argument(\"--rounds\", type=int, default=1)\n    parser.add_argument(\"--warmup\", action=\"store_true\")\n    parser.add_argument(\"--think\", default=None, help=\"true / false / low / medium / high, if supported by model and Ollama\")\n    args = parser.parse_args()\n\n    run_benchmark(args.models, args.prompt, args.rounds, args.warmup, args.think)\n```\n\nUsage:\n\n```bash\npython3 ollama_benchmark.py gemma4:26b qq36 --rounds 2 --warmup\npython3 ollama_benchmark.py qq36-nothink qq36 --think false --rounds 2 --warmup\n```\n\n### 6.4 Quality evaluation is required / 速度测试不能代替质量测试\n\nBefore using a model in a production workflow, test with real or sanitized materials:\n\n- Does it miss important facts?\n- Does it hallucinate?\n- Does it follow fixed output formats?\n- Can it handle long context?\n- Does it over-infer?\n- Does it leak visible reasoning when it should not?\n- Does it match the intended role?\n\n### 6.5 Real-Task Benchmark / 真实任务基准测试\n\nDo not test models only with one generic prompt. Ollama users usually care less about abstract speed and more about this question:\n\n> **Does this model perform well in my real task without breaking format, facts, or workflow stability?**\n\n建议为常用模型建立一组小型、固定、可复用的测试集。每次测试使用同一批任务、同一套评分维度、同一份结果记录模板。\n\nRecommended task set:\n\n1. **Short QA / 短问答**  \n   Tests basic understanding and daily response speed.\n\n2. **Long Summary / 长文摘要**  \n   Input 5,000–15,000 characters. Tests context handling, fact retention, and compression.\n\n3. **Structured Output / 结构化输出**  \n   Requires JSON, Markdown table, or fixed fields. Tests workflow reliability.\n\n4. **Role Task / 真实角色任务**  \n   Uses the user’s own real workflow task or a sanitized sample. Tests production fit.\n\nResult template:\n\n```markdown\n## Real-Task Benchmark Result\n\n- Date:\n- Tester:\n- Device:\n- Ollama version:\n- Model:\n- Quantization:\n- Think setting: true / false / low / medium / high / unset\n- Task type: Short QA / Long Summary / Structured Output / Role Task\n- prompt_eval_count:\n- eval_count:\n- load_duration:\n- prompt_eval_duration:\n- eval_duration:\n- tokens_per_second:\n- total_duration:\n- format_pass: yes / no\n- quality_score: 1–5\n- hallucination_risk: low / medium / high\n- production_ready: yes / no / observe\n- Notes:\n```\n\nQuality score:\n\n| Score | Meaning | Standard |\n|---:|---|---|\n| 5 | Production-ready | Acceptable speed, stable facts, stable format, not weaker than the current model |\n| 4 | Small-scope trial | Good quality but needs more observation |\n| 3 | Backup / testing only | Useful but has clear weaknesses |\n| 2 | Not recommended | Serious problems in speed, facts, or format |\n| 1 | Reject | Fails the task or frequently breaks output |\n\nKey rule:\n\n> A model should not be promoted into a production workflow only because it is fast. It should pass at least one real-task benchmark that matches the user’s actual use case.\n\n中文原则：\n\n> 不能因为一个模型跑得快，就把它放进正式工作流。至少要通过一个贴近真实使用场景的任务测试，才能替换生产模型。\n\nDecision guide:\n\n```text\nFast but unstable: keep as test only.\nHigh quality but slow: use as deep-analysis or backup model.\nFormat unstable: do not use for RAG, automation, or structured output.\nWins real-task benchmark: replace gradually and observe for 7–30 days.\n```\n\n---\n\n## 7. Aliases and Naming / 别名与命名\n\n### 7.1 Ollama model aliases / Ollama 模型别名\n\n```bash\nollama cp <source_model> <alias_name>\n```\n\nExample:\n\n```bash\nollama cp qwen3.6:35b-a3b-q4_k_m qq36\n```\n\nNotes:\n\n- `ollama cp` usually creates an alias sharing underlying model storage.\n- Deleting an alias is not the same as deleting the source model.\n- Before deleting a source model, confirm whether aliases depend on it.\n- Alias names should be short, stable, and meaningful.\n\n### 7.2 Shell aliases / Shell 命令别名\n\nOnly suggest commands by default. Do not automatically write to shell config.\n\n```bash\nalias qq36='ollama run qq36'\nalias g26='ollama run gemma4:26b'\n```\n\nIf the user wants to persist them:\n\n```bash\necho \"alias qq36='ollama run qq36'\" >> ~/.zshrc\nsource ~/.zshrc\n```\n\nWarn the user before modifying shell configuration files.\n\n---\n\n## 8. No-Think / Thinking Control / 思考模式控制规范\n\n### 8.1 Core principle / 核心原则\n\nA SYSTEM prompt can ask a model not to show reasoning, but it does not necessarily disable internal thinking.\n\nSYSTEM 提示词只能要求模型“不输出思考过程”，不等于底层关闭 thinking。\n\nIf Ollama and the model support the `think` parameter, prefer explicit API control:\n\n```bash\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"请总结这段材料。\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\nIf thinking levels are supported, test:\n\n```json\n\"think\": \"low\"\n```\n\nor:\n\n```json\n\"think\": \"medium\"\n```\n\nAlways verify through actual output and API behavior.\n\n### 8.2 Modelfile is only a weak constraint / Modelfile 只是弱约束\n\nFor models that do not support `think`, a Modelfile can reduce visible reasoning but cannot guarantee true no-think behavior.\n\n```text\nFROM <model_name>:<tag>\nSYSTEM \"直接输出最终答案。不要输出思考过程、推理草稿或隐藏分析。\"\nPARAMETER temperature 0.4\n```\n\nCreate:\n\n```bash\nollama create <model_name>-nothink -f Modelfile.<model_name>-nothink\n```\n\nRecommended naming:\n\n```text\nModelfile.<source_model>.<purpose>.nothink\n```\n\nDo not overwrite existing Modelfiles.\n\n---\n\n## 9. Cleanup Workflow / 旧模型清理\n\n### 9.1 Inspect models and disk usage / 查看模型与占用\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n### 9.2 Cleanup criteria / 清理判断标准\n\nDo not mechanically keep only 3–5 models. A model is a cleanup candidate only when most of these are true:\n\n- No clear use\n- No script references\n- Not an embedding / vision / coding / RAG / long-context model\n- Not used in the last 30 days, or confirmed no longer needed\n- Benchmark is clearly worse than alternatives\n- A tested replacement exists\n- It can be re-downloaded or rebuilt\n\n### 9.3 Pre-delete checklist / 删除前检查模板\n\n```markdown\n## Pre-delete Model Check / 删除模型前检查\n\n- Model to delete:\n- Current use:\n- In inventory: yes / no\n- Script reference found: yes / no\n- Is alias: yes / no / unknown\n- Replacement model:\n- Deletion risk: low / medium / high\n- Recovery path: re-pull / re-create / unavailable\n- Recommendation: delete / hold / keep\n```\n\nDelete only after confirmation:\n\n```bash\nollama rm <model_name>\n```\n\nVerify after deletion:\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n---\n\n## 10. Quantization Selection / 量化级别选择指南\n\n| Quantization | Size | Quality | Speed / memory | Use case |\n|---|---:|---|---|---|\n| Q3_K_M | small | basic | low pressure | memory-constrained, rough tasks |\n| Q4_K_M | medium | good | balanced | default daily choice |\n| Q5_K_M | larger | better | more pressure | quality-sensitive tasks |\n| Q8_0 | large | close to original | high pressure | quality-first and enough memory |\n\nRules:\n\n- Chat, summary, light workflows: start with Q4_K_M.\n- Higher-quality long-form or complex analysis: test Q5_K_M or Q8_0.\n- Low-memory devices: try smaller models or Q3_K_M.\n- Compare quantizations with the same benchmark set.\n\n---\n\n## 11. Troubleshooting / 常见故障\n\n### 11.1 Ollama version too old / Ollama 版本过旧\n\nSymptoms:\n\n```text\nunable to load model\n```\n\nPossible causes:\n\n- Ollama does not support the model format\n- CLI and server versions differ\n- Ollama server has not restarted\n\nCheck:\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\n```\n\nRestart or upgrade based on the installation method.\n\n### 11.2 Model download stuck / 模型下载卡住\n\nOrder of handling:\n\n1. Check network.\n2. Retry `ollama pull`; Ollama may resume.\n3. If official source remains stuck, consider ModelScope or Hugging Face.\n4. After using another source, validate model quality and format.\n\n### 11.3 Slow loading or memory pressure / 加载慢或内存不足\n\nCheck:\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\nSuggested actions:\n\n- Use a smaller model or lower quantization.\n- Reduce concurrency.\n- Avoid loading multiple large models at once.\n- Do not blindly raise `OLLAMA_CONTEXT_LENGTH`.\n- Use long context only when the task requires it.\n\n### 11.4 Output format unstable / 输出格式不稳定\n\nPossible causes:\n\n- Model weak at structured output\n- Temperature too high\n- Prompt too loose\n- Context too long and instructions diluted\n\nActions:\n\n- Lower temperature.\n- Use a stricter template.\n- Test with short structured samples.\n- Switch model if needed.\n\n---\n\n## 12. Recommended Workflows / 推荐工作流\n\n### 12.1 New model adoption / 新模型进入流程\n\n```text\nDiscover model → Check source → Define use case → Pull/import → Record environment → Benchmark → Real-task test → Add to inventory → Decide whether to bind to scripts\n```\n\n### 12.2 Model replacement / 模型替换流程\n\n```text\nTest candidate → Compare with current model → Check quality and format → Replace in small scope → Keep old model during observation → Update inventory → Decide whether to delete old model\n```\n\n### 12.3 Cleanup / 模型清理流程\n\n```text\nollama list → Inventory check → Script reference scan → Replacement check → Deletion recommendation → User confirmation → Delete → Verify disk usage → Update inventory\n```\n\n---\n\n## 13. Minimal Command Set / 最小可复制命令清单\n\n```bash\n# Environment\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n\n# Pull model\nollama pull <model_name>:<tag>\n\n# ModelScope fallback\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n\n# Create Ollama alias\nollama cp <source_model> <alias_name>\n\n# Scan script references\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n\n# Delete model, confirmation required\nollama rm <model_name>\n\n# API no-think test\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"请用一句话回答：测试。\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\n---\n\n## 14. Response Requirements / 输出要求\n\nWhen the user asks to manage Ollama models, respond with:\n\n1. Current judgment\n2. Safety risks\n3. Read-only checks to run first\n4. Whether confirmation is needed\n5. Next command or next decision\n\nWhen the user asks to clean up models, respond with:\n\n1. Model grouping\n2. Models that must not be deleted\n3. Models to observe\n4. Deletion candidates\n5. Script-reference scan command\n6. Explicit deletion confirmation prompt\n\nWhen the user asks to test models, respond with:\n\n1. Test goal\n2. Test scenarios\n3. Benchmark command\n4. Quality dimensions\n5. Production-readiness judgment\n\n---\n\n## 15. Version Notes / 版本说明\n\n### v1.4\n\n- Upgraded from the previous **v1.3 Model Pilot** release.\n- Renamed the maintained skill to **Ollama Model Pilot** while keeping `modelpilot` as the recommended package / slug for compatibility.\n- Added a clear v1.4 release highlight section near the top of the skill file.\n- Promoted **Real-Task Benchmark / 真实任务基准测试** into a core model-testing workflow, not just a supplementary note.\n- Added the production replacement rule: a model should not replace a workflow only because it is fast; it must pass at least one real-task benchmark matching the user’s actual use case.\n- Clarified legacy names: **Ollama Lifecycle Manager** and **Model Pilot**.\n- Clarified that old listings should be treated as legacy migration entries pointing to this maintained version.\n- Preserved the safety-first workflow: inventory first, script-reference scan before cleanup, confirmation before deletion, and cautious no-think handling.\n\n### v1.3\n\n- Released under the **Model Pilot** naming line.\n- Added / retained author note and safety-first positioning.\n- Included model inventory, script reference scanning, API benchmark, no-think rules, cleanup confirmation, and mirror-source risk warnings.\n\n---\n\n## 16. License / 许可说明\n\nIf this skill is derived from another published skill, confirm the original license and attribution requirements before distributing.\n\n如果本 skill 基于其他已发布 skill 改写或增强，请在发布前确认原始许可、署名和再分发要求。\n\nFile v1.4.0:README.md\n\n# Ollama Model Pilot v1.4\n\nA safety-first operational guide for managing local Ollama model workflows.\n\nThis v1.4 release is upgraded from the previous **v1.3 Model Pilot** release. The maintained name is now **Ollama Model Pilot**.\n\n## What it does\n\nOllama Model Pilot helps local model users manage:\n\n- model inventory and role tracking\n- safe cleanup with deletion confirmation\n- script reference scanning before removing models\n- reproducible API-based benchmark\n- **real-task benchmark for production model validation**\n- no-think / thinking-control usage\n- aliases and naming conventions\n- ModelScope / mirror fallback risk checks\n- common troubleshooting\n\n## Main v1.4 upgrade\n\nThe main upgrade in v1.4 is the **Real-Task Benchmark / 真实任务基准测试** workflow.\n\nThe skill now emphasizes that model testing should not stop at generic prompts or tokens-per-second. A model should be tested with fixed, repeatable tasks that match real usage:\n\n1. Short QA\n2. Long Summary\n3. Structured Output\n4. Role Task\n\nA model should not be promoted into a production workflow only because it is fast. It should pass at least one real-task benchmark that matches the user’s actual use case, with acceptable quality, fact stability, and output-format reliability.\n\n## Naming\n\nThis skill was previously published as **Model Pilot** and earlier as **Ollama Lifecycle Manager**. The maintained name is now:\n\n> **Ollama Model Pilot**\n\nThe package / slug can remain:\n\n```bash\nopenclaw skills install modelpilot\n```\n\nIf an older **Ollama Lifecycle Manager** listing still exists, it should be treated as a legacy migration entry and should point users to this maintained version.\n\n## Version\n\nCurrent version: **v1.4**\n\nPrevious release line: **v1.3 Model Pilot**\n\n## Files\n\n- `SKILL.md` — main skill file\n- `README.md` — publishing note\n- `CHANGELOG.md` — version history\n\nFile v1.4.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fj0qnbect2fmp59jj0vx34587hvns\",\n  \"slug\": \"modelpilot\",\n  \"version\": \"1.4.0\",\n  \"publishedAt\": 1779940753569\n}\n\nFile v1.4.0:CHANGELOG.md\n\n# Changelog\n\n## v1.4\n\n- Upgraded from the previous **v1.3 Model Pilot** release.\n- Renamed the maintained skill to **Ollama Model Pilot**.\n- Kept `modelpilot` as the recommended package / slug for compatibility.\n- Added a clear v1.4 release highlight section to the main `SKILL.md`.\n- Promoted **Real-Task Benchmark / 真实任务基准测试** into the core model-testing workflow.\n- Added the production replacement rule: do not replace a workflow only because a model is faster; require at least one matching real-task benchmark pass.\n- Clarified legacy names: **Ollama Lifecycle Manager** and **Model Pilot**.\n- Clarified that old listings should be used as legacy migration entries.\n- Preserved safety-first model management: inventory, script-reference scanning, deletion confirmation, no-think caution, and mirror-source risk checks.\n\n## v1.3\n\n- Released under the **Model Pilot** naming line.\n- Included author note and safety-first positioning.\n- Included model inventory, script reference scanning, API benchmark, no-think rules, cleanup confirmation, and mirror-source risk warnings.\n\nFile v1.4.0:skill-card.md\n\n## Description: <br>\nOllama Model Pilot is a safety-first operational guide for managing local Ollama model workflows, including model inventory, cleanup checks, benchmarking, aliases, and production-readiness validation. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[patmenciu](https://clawhub.ai/user/patmenciu) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and local LLM operators use this skill to inspect Ollama environments, maintain model inventories, benchmark real tasks, and make safer pull, alias, deletion, and production replacement decisions. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill may suggest commands that download, alias, create, or delete local Ollama models. <br>\nMitigation: Review each command before running it and require explicit confirmation for destructive operations such as model deletion. <br>\nRisk: Model cleanup can break scripts or workflows that reference a model name directly. <br>\nMitigation: Run the documented reference scan and inspect model inventory before approving removal or alias changes. <br>\nRisk: Benchmark results based only on speed can lead to unsuitable production model replacement. <br>\nMitigation: Use repeatable real-task benchmarks and check quality, fact stability, and output-format reliability before promotion. <br>\nRisk: Persistent shell or Ollama configuration changes can affect future local model behavior. <br>\nMitigation: Only approve shell configuration or environment changes when the intended impact is clear. <br>\n\n\n## Reference(s): <br>\n- [ClawHub listing](https://clawhub.ai/patmenciu/modelpilot) <br>\n- [README.md](README.md) <br>\n- [CHANGELOG.md](CHANGELOG.md) <br>\n- [SKILL.md](SKILL.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [guidance, markdown, code, shell commands, configuration] <br>\n**Output Format:** [Markdown guidance with checklists, command examples, tables, and code snippets] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Includes human confirmation steps before destructive or persistent Ollama changes.] <br>\n\n## Skill Version(s): <br>\n1.4.0 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.3.0: 4 files, 17922 bytes\n\nFiles: skill-card.md (2593b), SKILL.en.md (20714b), SKILL.md (19383b), _meta.json (129b)\n\nFile v1.3.0:SKILL.md\n\n---\nname: Model Pilot\ndescription: \"Your copilot for local LLM management — safely pull, benchmark, alias, and clean up models with confidence.\"\ndescription_zh: \"本地大模型管理副驾驶——安全拉取、测速、别名、清理，全程可控无忧。\"\nagent_created: true\ntags: [ollama, llm, local-models, benchmark, model-management, safety, no-think]\nsource: \"https://clawhub.ai/skills/model-pilot\"\n---\n\n# Model Pilot — 本地大模型管理技能（安全增强版）\n\n> **语言**：中文（简体） | English version: `SKILL.en.md`\n\n## 用途\n\n本 skill 是 Ollama 本地模型生命周期管理的**操作指南与检查清单**，用于帮助用户安全地完成：\n\n- 拉取新模型\n- 记录模型台账\n- 进行可复盘 Benchmark\n- 创建模型别名\n- 创建 no-think / low-think 使用方式\n- 扫描脚本中的模型引用\n- 清理旧模型\n- 处理国内镜像回退和常见故障\n\n触发词：Ollama、拉模型、下载模型、测速、Benchmark、模型对比、模型别名、清理模型、ModelScope、GGUF、no-think、think false、模型台账、模型生命周期。\n\n---\n\n## 0. 安全边界\n\n本 skill 默认遵循\"只读优先、确认后修改、破坏性操作单独确认\"的原则。\n\n### 默认允许的只读操作\n\n可以直接执行或建议执行：\n\n```bash\nollama list\nollama --version\ncurl http://localhost:11434/api/version\ncurl http://localhost:11434/api/tags\npwd\nls\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\ndu -sh ~/.ollama/models/\n```\n\n### 必须用户确认后才能执行的操作\n\n以下操作涉及下载、复制、创建、删除或修改配置，必须先说明影响，再等待用户确认：\n\n```bash\nollama pull <model>\nollama cp <source_model> <alias>\nollama create <new_model> -f <Modelfile>\nollama rm <model>\nbrew upgrade ollama\n```\n\n### 不应自动执行的操作\n\n除非用户明确要求，不要自动执行：\n\n- 删除模型：`ollama rm`\n- 覆盖已有 Modelfile\n- 修改 `~/.zshrc`、`~/.bashrc`、`~/.profile`\n- 修改系统级环境变量\n- 改动 `OLLAMA_CONTEXT_LENGTH`、`OLLAMA_NUM_PARALLEL`、`OLLAMA_MAX_LOADED_MODELS`\n- 批量下载多个大模型\n\n### 删除模型前必须输出确认信息\n\n删除任何模型前，必须先输出：\n\n1. 当前 `ollama list` 摘要\n2. 待删除模型名称\n3. 模型用途判断\n4. 是否发现脚本引用\n5. 是否存在替代模型\n6. 删除后能否重新拉取\n7. 用户确认语句\n\n推荐确认语句：\n\n```text\n请确认是否删除模型：<model_name>。删除后如需恢复，需要重新下载或重新 create。\n```\n\n---\n\n## 1. 适用场景与不适用场景\n\n### 适用场景\n\n- 本地模型越来越多，需要建立台账\n- 想比较多个模型在同一设备上的速度和质量\n- 想创建简短别名，方便命令行调用\n- 想判断某个模型是否可以安全删除\n- 想处理 Ollama 官方 registry 下载慢、卡住、失败\n- 想为不同工作流选择不同模型\n- 想规范 no-think / think false 使用方式\n\n### 不适用场景\n\n- 自动替用户决定删除哪些模型\n- 自动替用户选择\"最强模型\"\n- 自动修改生产脚本中的模型名称\n- 自动处理涉及隐私材料的 Benchmark 输入\n- 代替模型质量评估或专业任务评审\n\n---\n\n## 2. 环境检查\n\n在做任何模型管理前，先检查环境。\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n```\n\n建议记录：\n\n```markdown\n## Ollama 环境记录\n\n- 日期：YYYY-MM-DD\n- 操作系统：\n- 芯片/CPU/GPU：\n- 内存：\n- Ollama CLI 版本：\n- Ollama Server 版本：\n- 模型目录占用：\n- 当前主要用途：聊天 / RAG / 长文分析 / 代码 / 视觉 / 其他\n```\n\n如果 CLI 和 server 版本不一致，优先重启 Ollama 服务后再测试。\n\n---\n\n## 3. 模型台账\n\n生命周期管理的核心不是\"模型多就删\"，而是知道每个模型为什么存在。\n\n建议维护一个 `Ollama模型台账.md`。\n\n```markdown\n# Ollama 模型台账\n\n| 模型名 | 类型 | 当前用途 | 绑定脚本/应用 | 状态 | 是否可删 | 替代模型 | 最近测试日期 | 备注 |\n|---|---|---|---|---|---|---|---|---|\n| gemma4:26b | 通用文本 | 长文分析 | batch_xxx.py | 核心 | 否 | - | 2026-xx-xx | 稳定 |\n| qwen3-vl | 视觉模型 | 图片理解 | photo_app | 核心 | 否 | - | 2026-xx-xx | 保留 |\n| gpt-oss:20b | 通用文本 | 测试 | 无 | 观察 | 可考虑 | xxx | 2026-xx-xx | 速度慢 |\n```\n\n### 状态分层\n\n建议把模型分为五类：\n\n1. **核心模型**：默认工作流正在使用，不能删除。\n2. **任务模型**：embedding、vision、coding、long-context 等特殊用途模型。\n3. **备用模型**：核心模型失效或质量不足时备用。\n4. **测试模型**：新模型观察期，建议 7—30 天后复盘。\n5. **废弃模型**：无明确用途、无脚本引用、已有替代，可考虑删除。\n\n---\n\n## 4. 脚本引用扫描\n\n清理模型前，先扫描当前项目是否有脚本引用模型名称。\n\n在项目根目录执行：\n\n```bash\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n```\n\n如果没有 `rg`，可用：\n\n```bash\ngrep -R \"ollama run\\|MODEL_NAME\\|model_name\\|\\\"model\\\"\" . 2>/dev/null\n```\n\n输出后整理成：\n\n```markdown\n## 模型引用扫描结果\n\n| 模型名 | 文件 | 行号/位置 | 用途 | 是否影响删除 |\n|---|---|---|---|---|\n| gemma-doc-nothink | batch_qwen_cards.py | MODEL_NAME | 资料抽取 | 是 |\n| mistral-qa | qa_ask.py | model_name | 深度问答 | 是 |\n```\n\n删除模型前必须回答：\n\n1. 是否被脚本硬编码引用？\n2. 是否是 alias 的源模型？\n3. 是否用于 embedding / RAG / 视觉 / 代码 / 长上下文？\n4. 是否已有质量和速度都更好的替代模型？\n5. 删除后是否容易恢复？\n\n---\n\n## 5. 拉取新模型流程\n\n### 5.1 官方 registry 优先\n\n```bash\nollama pull <model_name>:<tag>\n```\n\n拉取前先确认：\n\n- 模型来源\n- 模型尺寸\n- 量化等级\n- 预计磁盘占用\n- 本机内存是否足够\n- 是否已有同类模型\n- 是否有明确测试任务\n\n### 5.2 下载失败后的回退流程\n\n如果官方 registry 卡住或超时：\n\n1. 先等待一段合理时间，确认不是短暂网络波动。\n2. 终止卡住的下载。\n3. 到 ModelScope 或 Hugging Face 查找对应模型。\n4. 核对发布者、文件名、量化格式、更新时间、license。\n5. 使用完整路径重新拉取或手动导入。\n\nModelScope 示例：\n\n```bash\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n```\n\n### 5.3 ModelScope 风险提示\n\n使用镜像源时必须注意：\n\n- 路径和 tag 以模型页面为准。\n- 同名模型可能不是官方发布者上传。\n- template、chat format、license、量化格式可能与官方版本不同。\n- 拉取后必须用固定测试集验证输出格式、速度和质量。\n- 工作流模型不要只因下载成功就直接替换生产模型。\n\n---\n\n## 6. Benchmark 标准流程\n\n不要只用 `ollama run + subprocess` 统计总耗时。总耗时会混入模型加载、命令行包装和输出渲染，不适合严肃对比。\n\n优先调用 Ollama API，读取：\n\n- `total_duration`\n- `load_duration`\n- `prompt_eval_count`\n- `prompt_eval_duration`\n- `eval_count`\n- `eval_duration`\n\n核心速度指标：\n\n```text\ntokens_per_second = eval_count / (eval_duration / 1e9)\n```\n\n### 6.1 Benchmark 场景\n\n建议至少分四类测试：\n\n1. **短问答**：100—300 字输出，测试日常响应。\n2. **长文写作**：800—1500 字输出，测试持续生成。\n3. **长文档摘要**：输入 5000—15000 字，测试上下文处理。\n4. **结构化输出**：要求 JSON 或固定 Markdown，测试工作流稳定性。\n\n### 6.2 冷启动与热启动\n\n- 冷启动：包含模型加载时间，用于判断首次响应体验。\n- 热启动：先预热一次，再测试，用于判断真实连续使用速度。\n\n### 6.3 API Benchmark 脚本\n\n保存为 `ollama_benchmark.py`：\n\n```python\n#!/usr/bin/env python3\nimport argparse\nimport json\nimport statistics\nimport time\nimport urllib.request\nfrom typing import Any, Dict, List\n\nOLLAMA_URL = \"http://localhost:11434/api/generate\"\n\n\ndef call_ollama(model: str, prompt: str, think: str | None = None, timeout: int = 600) -> Dict[str, Any]:\n    payload: Dict[str, Any] = {\n        \"model\": model,\n        \"prompt\": prompt,\n        \"stream\": False,\n    }\n\n    if think is not None:\n        if think.lower() == \"false\":\n            payload[\"think\"] = False\n        elif think.lower() == \"true\":\n            payload[\"think\"] = True\n        else:\n            payload[\"think\"] = think\n\n    req = urllib.request.Request(\n        OLLAMA_URL,\n        data=json.dumps(payload).encode(\"utf-8\"),\n        headers={\"Content-Type\": \"application/json\"},\n        method=\"POST\",\n    )\n\n    start = time.time()\n    with urllib.request.urlopen(req, timeout=timeout) as resp:\n        data = json.loads(resp.read().decode(\"utf-8\"))\n    wall_sec = time.time() - start\n\n    eval_count = data.get(\"eval_count\") or 0\n    eval_duration = data.get(\"eval_duration\") or 0\n    prompt_eval_count = data.get(\"prompt_eval_count\") or 0\n    prompt_eval_duration = data.get(\"prompt_eval_duration\") or 0\n    load_duration = data.get(\"load_duration\") or 0\n    total_duration = data.get(\"total_duration\") or 0\n    response = data.get(\"response\", \"\")\n\n    tps = None\n    if eval_count and eval_duration:\n        tps = eval_count / (eval_duration / 1e9)\n\n    prompt_tps = None\n    if prompt_eval_count and prompt_eval_duration:\n        prompt_tps = prompt_eval_count / (prompt_eval_duration / 1e9)\n\n    return {\n        \"model\": model,\n        \"wall_sec\": round(wall_sec, 3),\n        \"total_sec_api\": round(total_duration / 1e9, 3) if total_duration else None,\n        \"load_sec\": round(load_duration / 1e9, 3) if load_duration else None,\n        \"prompt_eval_count\": prompt_eval_count,\n        \"prompt_eval_sec\": round(prompt_eval_duration / 1e9, 3) if prompt_eval_duration else None,\n        \"prompt_tokens_per_sec\": round(prompt_tps, 2) if prompt_tps else None,\n        \"eval_count\": eval_count,\n        \"eval_sec\": round(eval_duration / 1e9, 3) if eval_duration else None,\n        \"tokens_per_sec\": round(tps, 2) if tps else None,\n        \"response_chars\": len(response),\n        \"preview\": response[:160].replace(\"\\n\", \" \") + (\"...\" if len(response) > 160 else \"\"),\n    }\n\n\ndef run_benchmark(models: List[str], prompt: str, rounds: int, warmup: bool, think: str | None) -> None:\n    if warmup:\n        for model in models:\n            try:\n                call_ollama(model, \"请用一句话回答：测试。\", think=think)\n            except Exception as e:\n                print(f\"[WARMUP FAIL] {model}: {e}\")\n\n    for model in models:\n        results = []\n        for i in range(rounds):\n            try:\n                result = call_ollama(model, prompt, think=think)\n                results.append(result)\n                print(json.dumps(result, ensure_ascii=False))\n            except Exception as e:\n                print(json.dumps({\"model\": model, \"round\": i + 1, \"error\": str(e)}, ensure_ascii=False))\n\n        speeds = [r[\"tokens_per_sec\"] for r in results if r.get(\"tokens_per_sec\") is not None]\n        if speeds:\n            summary = {\n                \"model\": model,\n                \"rounds\": len(speeds),\n                \"tokens_per_sec_avg\": round(statistics.mean(speeds), 2),\n                \"tokens_per_sec_min\": round(min(speeds), 2),\n                \"tokens_per_sec_max\": round(max(speeds), 2),\n            }\n            print(\"[SUMMARY] \" + json.dumps(summary, ensure_ascii=False))\n\n\nif __name__ == \"__main__\":\n    parser = argparse.ArgumentParser(description=\"Ollama API benchmark\")\n    parser.add_argument(\"models\", nargs=\"+\", help=\"Models to benchmark\")\n    parser.add_argument(\"--prompt\", default=\"请写一段约500字的中文说明，介绍本地大模型的实际用途。\")\n    parser.add_argument(\"--rounds\", type=int, default=1)\n    parser.add_argument(\"--warmup\", action=\"store_true\")\n    parser.add_argument(\"--think\", default=None, help=\"true / false / low / medium / high, if supported by model and Ollama\")\n    args = parser.parse_args()\n\n    run_benchmark(args.models, args.prompt, args.rounds, args.warmup, args.think)\n```\n\n使用示例：\n\n```bash\npython3 ollama_benchmark.py gemma4:26b qq36 --rounds 2 --warmup\npython3 ollama_benchmark.py qq36-nothink qq36 --think false --rounds 2 --warmup\n```\n\n### 6.4 质量评估必须用真实任务\n\n速度测试不能代替质量测试。模型用于生产工作流前，必须用真实材料或脱敏材料测试：\n\n- 是否漏事实\n- 是否编造\n- 是否能按固定格式输出\n- 是否能处理长文本\n- 是否容易过度推断\n- 是否会输出思考过程\n- 是否符合任务角色定位\n\n---\n\n## 7. 创建模型别名\n\n### 7.1 Ollama 模型别名\n\n```bash\nollama cp <source_model> <alias_name>\n```\n\n示例：\n\n```bash\nollama cp qwen3.6:35b-a3b-q4_k_m qq36\n```\n\n注意：\n\n- `ollama cp` 创建的别名通常共享底层模型存储。\n- 删除 alias 不等于删除源模型。\n- 删除源模型前要确认 alias 是否仍可用。\n- alias 名称应短、稳定、能反映模型家族或任务用途。\n\n### 7.2 Shell 命令别名\n\n只建议输出命令，不自动写入 shell 配置。\n\n```bash\nalias qq36='ollama run qq36'\nalias g26='ollama run gemma4:26b'\n```\n\n如需写入：\n\n```bash\necho \"alias qq36='ollama run qq36'\" >> ~/.zshrc\nsource ~/.zshrc\n```\n\n写入前必须提醒用户：这会修改 shell 配置文件。\n\n---\n\n## 8. no-think / thinking 控制规范\n\n### 8.1 关键原则\n\nSYSTEM 提示词只能要求模型\"不输出思考过程\"，不等于底层关闭 thinking。\n\n如果 Ollama 和模型支持 `think` 参数，应优先用 API 显式设置：\n\n```bash\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"请总结这段材料。\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\n如果模型支持 thinking levels，也可以测试：\n\n```json\n\"think\": \"low\"\n```\n\n或：\n\n```json\n\"think\": \"medium\"\n```\n\n具体是否生效，必须以模型实际输出和 Ollama API 支持情况为准。\n\n### 8.2 Modelfile 方式只是弱约束\n\n对不支持 `think` 参数的模型，可用 Modelfile 提示减少显式思考，但不能保证真正关闭 thinking。\n\n```text\nFROM <model_name>:<tag>\nSYSTEM \"直接输出最终答案。不要输出思考过程、推理草稿或隐藏分析。\"\nPARAMETER temperature 0.4\n```\n\n创建：\n\n```bash\nollama create <model_name>-nothink -f Modelfile.<model_name>-nothink\n```\n\n建议文件命名：\n\n```text\nModelfile.<source_model>.<purpose>.nothink\n```\n\n不要覆盖已有 `Modelfile`。\n\n---\n\n## 9. 清理旧模型\n\n### 9.1 查看模型与占用\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n### 9.2 清理判断标准\n\n不要机械执行\"只保留 3—5 个模型\"。更合理的判断是：\n\n- 无明确用途\n- 无脚本引用\n- 非 embedding / vision / coding / RAG 等特殊任务模型\n- 最近 30 天未使用或已确认不再使用\n- benchmark 明显落后\n- 已有同类替代模型\n- 删除后可重新下载或重建\n\n### 9.3 删除前检查模板\n\n```markdown\n## 删除模型前检查\n\n- 待删除模型：\n- 当前用途：\n- 是否在台账中：是 / 否\n- 是否发现脚本引用：是 / 否\n- 是否是 alias：是 / 否\n- 是否有替代模型：\n- 删除风险：低 / 中 / 高\n- 恢复方式：重新 pull / 重新 create / 无法恢复\n- 建议：删除 / 暂缓 / 保留\n```\n\n删除命令：\n\n```bash\nollama rm <model_name>\n```\n\n删除后确认：\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n---\n\n## 10. 量化级别选择指南\n\n| 量化 | 体积 | 质量 | 速度/内存 | 适用场景 |\n|---|---:|---|---|---|\n| Q3_K_M | 小 | 一般 | 压力小 | 内存紧张、粗略任务 |\n| Q4_K_M | 中 | 良好 | 平衡 | 日常首选 |\n| Q5_K_M | 较大 | 较好 | 压力略高 | 对质量更敏感 |\n| Q8_0 | 大 | 接近原始 | 压力大 | 质量优先且内存充足 |\n\n选择原则：\n\n- 普通聊天、摘要、轻量工作流：优先 Q4_K_M。\n- 高质量长文、复杂分析：可测试 Q5_K_M 或 Q8_0。\n- 小内存设备：先试 Q3_K_M 或更小模型。\n- 不同量化版本必须用同一测试集比较，不要凭感觉判断。\n\n---\n\n## 11. 常见故障\n\n### 11.1 Ollama 版本过旧导致模型加载失败\n\n症状：\n\n```text\nunable to load model\n```\n\n可能原因：\n\n- Ollama 版本不支持新版 GGUF\n- CLI 和 server 版本不一致\n- server 未重启\n\n处理：\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\n```\n\n然后根据安装方式升级或重启 Ollama。\n\n### 11.2 模型下载卡住\n\n处理顺序：\n\n1. 确认网络是否正常。\n2. 重新执行 `ollama pull`，Ollama 通常可断点续传。\n3. 官方源长期卡住时，考虑 ModelScope 或 Hugging Face。\n4. 换源后必须重新验证模型质量和格式。\n\n### 11.3 模型加载慢或内存不足\n\n检查：\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n建议：\n\n- 降低模型尺寸或量化等级。\n- 减少并发。\n- 避免同时驻留多个大模型。\n- 不要盲目提高 `OLLAMA_CONTEXT_LENGTH`。\n- 长上下文只在任务确实需要时开启。\n\n### 11.4 输出格式不稳定\n\n可能原因：\n\n- 模型不适合结构化输出\n- temperature 过高\n- prompt 太松\n- 上下文过长导致指令稀释\n\n处理：\n\n- 降低 temperature。\n- 增加固定模板。\n- 用短样本测试格式遵循能力。\n- 必要时更换模型。\n\n---\n\n## 12. 推荐工作流\n\n### 新模型进入流程\n\n```text\n发现模型 → 检查来源 → 判断用途 → 拉取/导入 → 环境记录 → Benchmark → 真实任务测试 → 写入台账 → 决定是否绑定脚本\n```\n\n### 模型替换流程\n\n```text\n候选模型测试 → 与现有模型对比 → 检查输出质量 → 小范围替换 → 保留旧模型一段观察期 → 稳定后更新台账 → 再决定是否删除旧模型\n```\n\n### 模型清理流程\n\n```text\nollama list → 台账核对 → 脚本引用扫描 → 判断替代关系 → 输出删除建议 → 用户确认 → 删除 → 复查空间 → 更新台账\n```\n\n---\n\n## 13. 最小可复制命令清单\n\n```bash\n# 环境\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n\n# 拉模型\nollama pull <model_name>:<tag>\n\n# ModelScope 回退\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n\n# 创建 Ollama alias\nollama cp <source_model> <alias_name>\n\n# 扫描脚本引用\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n\n# 删除模型，必须确认后执行\nollama rm <model_name>\n\n# API no-think 测试\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"请用一句话回答：测试。\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\n---\n\n## 14. 输出要求\n\n当用户请求\"帮我管理 Ollama 模型\"时，应优先输出：\n\n1. 当前判断\n2. 安全风险\n3. 建议执行的只读检查\n4. 是否需要用户确认\n5. 下一步命令\n\n当用户请求\"清理模型\"时，应输出：\n\n1. 模型分层\n2. 不可删模型\n3. 可观察模型\n4. 可删除候选\n5. 删除前引用扫描命令\n6. 删除确认提示\n\n当用户请求\"测试模型\"时，应输出：\n\n1. 测试目标\n2. 测试场景\n3. Benchmark 命令\n4. 质量评估维度\n5. 是否建议进入生产工作流\n\nFile v1.3.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fj0qnbect2fmp59jj0vx34587hvns\",\n  \"slug\": \"modelpilot\",\n  \"version\": \"1.3.0\",\n  \"publishedAt\": 1779876754936\n}\n\nFile v1.3.0:skill-card.md\n\n## Description: <br>\nYour copilot for local LLM management - safely pull, benchmark, alias, and clean up models with confidence. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[patmenciu](https://clawhub.ai/user/patmenciu) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nExternal developers, engineers, and local LLM users use Model Pilot to manage Ollama model lifecycles: pull and benchmark models, maintain registries, create aliases and no-think variants, scan references, and clean up with confirmation gates. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Downloads, deletions, upgrades, and shell configuration edits can affect disk usage or the local Ollama setup. <br>\nMitigation: Review proposed commands and require explicit user confirmation before model pulls, removals, alias creation, upgrades, or shell file writes. <br>\nRisk: Mirror or fallback model sources may differ from official sources in publisher, template, license, quantization, or chat format. <br>\nMitigation: Verify source metadata and validate output format, speed, and quality with a fixed test set before replacing workflow models. <br>\nRisk: Benchmark results can be misleading if they measure only speed or include private materials. <br>\nMitigation: Use anonymized or task-appropriate prompts and combine speed metrics with real-task quality checks. <br>\nRisk: A Modelfile prompt may reduce visible thinking output but cannot guarantee thinking is disabled. <br>\nMitigation: Prefer Ollama's supported think parameter when available and validate model behavior before binding it to a workflow. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/patmenciu/modelpilot) <br>\n- [Publisher profile](https://clawhub.ai/user/patmenciu) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [guidance, markdown, code, shell commands, configuration] <br>\n**Output Format:** [Markdown with inline bash, Python, JSON, and configuration snippets] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Includes confirmation checkpoints for downloads, deletions, aliases, upgrades, and shell configuration changes.] <br>\n\n## Skill Version(s): <br>\n1.3.0 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v1.3.0:SKILL.en.md\n\n---\nname: Model Pilot\ndescription: \"Your copilot for local LLM management — safely pull, benchmark, alias, and clean up models with confidence.\"\ndescription_zh: \"本地大模型管理副驾驶——安全拉取、测速、别名、清理，全程可控无忧。\"\nagent_created: true\ntags: [ollama, llm, local-models, benchmark, model-management, safety, no-think]\nsource: \"https://clawhub.ai/skills/model-pilot\"\nlang: en\n---\n\n# Model Pilot — Local LLM Management Skill (Safety-Enhanced Edition)\n\n> **Language**: English | 中文版本：`SKILL.md`\n\n## Purpose\n\nThis skill is an **operation guide and checklist** for managing the full lifecycle of local Ollama models, helping users safely:\n\n- Pull new models\n- Maintain a model registry\n- Run reproducible benchmarks\n- Create model aliases\n- Set up no-think / low-think modes\n- Scan scripts for model references\n- Clean up old models\n- Handle mirror fallbacks and common issues\n\nTrigger words: Ollama, pull model, download model, benchmark, model comparison, model alias, cleanup model, ModelScope, GGUF, no-think, think false, model registry, model lifecycle.\n\n---\n\n## 0. Safety Boundaries\n\nThis skill follows the principle of \"read-first, confirm-before-modify, separate confirmation for destructive operations.\"\n\n### Read-Only Operations (Auto-Allow)\n\nCan be executed or suggested directly:\n\n```bash\nollama list\nollama --version\ncurl http://localhost:11434/api/version\ncurl http://localhost:11434/api/tags\npwd\nls\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\ndu -sh ~/.ollama/models/\n```\n\n### Operations Requiring User Confirmation\n\nThe following involve downloading, copying, creating, deleting, or modifying configuration. Explain the impact first, then wait for confirmation:\n\n```bash\nollama pull <model>\nollama cp <source_model> <alias>\nollama create <new_model> -f <Modelfile>\nollama rm <model>\nbrew upgrade ollama\n```\n\n### Operations That Should Not Run Automatically\n\nUnless the user explicitly requests, do not:\n\n- Delete models: `ollama rm`\n- Overwrite existing Modelfiles\n- Modify `~/.zshrc`, `~/.bashrc`, `~/.profile`\n- Modify system-level environment variables\n- Change `OLLAMA_CONTEXT_LENGTH`, `OLLAMA_NUM_PARALLEL`, `OLLAMA_MAX_LOADED_MODELS`\n- Batch-download multiple large models\n\n### Pre-Deletion Confirmation\n\nBefore deleting any model, output:\n\n1. Current `ollama list` summary\n2. Model name to be deleted\n3. Model purpose assessment\n4. Whether script references were found\n5. Whether alternative models exist\n6. Whether the model can be re-pulled\n7. User confirmation statement\n\nRecommended confirmation statement:\n\n```text\nPlease confirm deletion of model: <model_name>. To restore after deletion, you will need to re-download or re-create it.\n```\n\n---\n\n## 1. Scope\n\n### In Scope\n\n- Local model collection growing and needing a registry\n- Comparing multiple models on the same hardware\n- Creating short aliases for command-line convenience\n- Determining whether a model can be safely deleted\n- Handling slow, stuck, or failed downloads from the Ollama registry\n- Selecting models for different workflows\n- Standardizing no-think / think false usage\n\n### Out of Scope\n\n- Automatically deciding which models to delete\n- Automatically choosing the \"best model\"\n- Automatically modifying model names in production scripts\n- Handling benchmark inputs involving private materials\n- Replacing professional model quality assessments\n\n---\n\n## 2. Environment Check\n\nBefore any model management, check the environment first.\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n```\n\nRecommended record:\n\n```markdown\n## Ollama Environment Record\n\n- Date: YYYY-MM-DD\n- OS:\n- Chip/CPU/GPU:\n- Memory:\n- Ollama CLI version:\n- Ollama Server version:\n- Model directory size:\n- Primary use case: Chat / RAG / Long-form analysis / Code / Vision / Other\n```\n\nIf CLI and server versions mismatch, restart the Ollama service before testing.\n\n---\n\n## 3. Model Registry\n\nThe core of lifecycle management is not \"delete when you have too many models\" — it's knowing why each model exists.\n\nConsider maintaining an `ollama-model-registry.md`.\n\n```markdown\n# Ollama Model Registry\n\n| Model | Type | Current Use | Bound Scripts/Apps | Status | Deletable | Replacement | Last Tested | Notes |\n|---|---|---|---|---|---|---|---|---|\n| gemma4:26b | General text | Long-form analysis | batch_xxx.py | Core | No | - | 2026-xx-xx | Stable |\n| qwen3-vl | Vision | Image understanding | photo_app | Core | No | - | 2026-xx-xx | Keep |\n| gpt-oss:20b | General text | Testing | None | Trial | Consider | xxx | 2026-xx-xx | Slow |\n```\n\n### Status Tiers\n\nConsider categorizing models into five tiers:\n\n1. **Core models**: Actively used in default workflows. Must not be deleted.\n2. **Task models**: Special-purpose models for embedding, vision, coding, long-context, etc.\n3. **Backup models**: Fallbacks when core models fail or fall short on quality.\n4. **Trial models**: New models under observation. Review after 7–30 days.\n5. **Deprecated models**: No clear purpose, no script references, already replaced. Candidates for deletion.\n\n---\n\n## 4. Script Reference Scan\n\nBefore cleaning up models, scan the current project for scripts referencing model names.\n\nRun in the project root:\n\n```bash\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n```\n\nIf `rg` is not available:\n\n```bash\ngrep -R \"ollama run\\|MODEL_NAME\\|model_name\\|\\\"model\\\"\" . 2>/dev/null\n```\n\nOrganize results into:\n\n```markdown\n## Model Reference Scan Results\n\n| Model | File | Line/Position | Purpose | Affects Deletion |\n|---|---|---|---|---|\n| gemma-doc-nothink | batch_qwen_cards.py | MODEL_NAME | Data extraction | Yes |\n| mistral-qa | qa_ask.py | model_name | Deep Q&A | Yes |\n```\n\nBefore deleting a model, answer:\n\n1. Is it hardcoded in any script?\n2. Is it the source model for an alias?\n3. Is it used for embedding / RAG / vision / code / long-context?\n4. Is there a better alternative in both quality and speed?\n5. Can it be easily recovered after deletion?\n\n---\n\n## 5. Pulling New Models\n\n### 5.1 Official Registry First\n\n```bash\nollama pull <model_name>:<tag>\n```\n\nBefore pulling, confirm:\n\n- Model source\n- Model size\n- Quantization level\n- Expected disk usage\n- Whether local memory is sufficient\n- Whether a similar model already exists\n- Whether there is a clear testing task\n\n### 5.2 Fallback Flow on Download Failure\n\nIf the official registry is stuck or times out:\n\n1. Wait a reasonable amount of time to confirm it's not a transient network issue.\n2. Terminate the stuck download.\n3. Search ModelScope or Hugging Face for the corresponding model.\n4. Verify the publisher, filename, quantization format, update date, and license.\n5. Re-pull or manually import using the full path.\n\nModelScope example:\n\n```bash\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n```\n\n### 5.3 ModelScope Risk Warnings\n\nWhen using mirror sources:\n\n- Path and tag should match the model's page exactly.\n- Models with the same name may not be uploaded by the official publisher.\n- Template, chat format, license, and quantization format may differ from the official version.\n- After pulling, validate output format, speed, and quality with a fixed test set.\n- Do not replace production models solely because a mirror download succeeded.\n\n---\n\n## 6. Benchmark Standard Workflow\n\nDo not rely solely on `ollama run + subprocess` for total elapsed time. Total time includes model loading, CLI wrapping, and output rendering — it is not suitable for rigorous comparison.\n\nPrefer calling the Ollama API and reading:\n\n- `total_duration`\n- `load_duration`\n- `prompt_eval_count`\n- `prompt_eval_duration`\n- `eval_count`\n- `eval_duration`\n\nCore speed metric:\n\n```text\ntokens_per_second = eval_count / (eval_duration / 1e9)\n```\n\n### 6.1 Benchmark Scenarios\n\nRecommend at least four test categories:\n\n1. **Short Q&A**: 100–300 word output, testing daily responsiveness.\n2. **Long-form writing**: 800–1500 word output, testing sustained generation.\n3. **Long document summarization**: 5,000–15,000 word input, testing context processing.\n4. **Structured output**: Requesting JSON or fixed Markdown, testing workflow stability.\n\n### 6.2 Cold Start vs. Warm Start\n\n- Cold start: Includes model loading time. For evaluating first-response experience.\n- Warm start: Warm up once, then test. For evaluating real sustained-use speed.\n\n### 6.3 API Benchmark Script\n\nSave as `ollama_benchmark.py`:\n\n```python\n#!/usr/bin/env python3\nimport argparse\nimport json\nimport statistics\nimport time\nimport urllib.request\nfrom typing import Any, Dict, List\n\nOLLAMA_URL = \"http://localhost:11434/api/generate\"\n\n\ndef call_ollama(model: str, prompt: str, think: str | None = None, timeout: int = 600) -> Dict[str, Any]:\n    payload: Dict[str, Any] = {\n        \"model\": model,\n        \"prompt\": prompt,\n        \"stream\": False,\n    }\n\n    if think is not None:\n        if think.lower() == \"false\":\n            payload[\"think\"] = False\n        elif think.lower() == \"true\":\n            payload[\"think\"] = True\n        else:\n            payload[\"think\"] = think\n\n    req = urllib.request.Request(\n        OLLAMA_URL,\n        data=json.dumps(payload).encode(\"utf-8\"),\n        headers={\"Content-Type\": \"application/json\"},\n        method=\"POST\",\n    )\n\n    start = time.time()\n    with urllib.request.urlopen(req, timeout=timeout) as resp:\n        data = json.loads(resp.read().decode(\"utf-8\"))\n    wall_sec = time.time() - start\n\n    eval_count = data.get(\"eval_count\") or 0\n    eval_duration = data.get(\"eval_duration\") or 0\n    prompt_eval_count = data.get(\"prompt_eval_count\") or 0\n    prompt_eval_duration = data.get(\"prompt_eval_duration\") or 0\n    load_duration = data.get(\"load_duration\") or 0\n    total_duration = data.get(\"total_duration\") or 0\n    response = data.get(\"response\", \"\")\n\n    tps = None\n    if eval_count and eval_duration:\n        tps = eval_count / (eval_duration / 1e9)\n\n    prompt_tps = None\n    if prompt_eval_count and prompt_eval_duration:\n        prompt_tps = prompt_eval_count / (prompt_eval_duration / 1e9)\n\n    return {\n        \"model\": model,\n        \"wall_sec\": round(wall_sec, 3),\n        \"total_sec_api\": round(total_duration / 1e9, 3) if total_duration else None,\n        \"load_sec\": round(load_duration / 1e9, 3) if load_duration else None,\n        \"prompt_eval_count\": prompt_eval_count,\n        \"prompt_eval_sec\": round(prompt_eval_duration / 1e9, 3) if prompt_eval_duration else None,\n        \"prompt_tokens_per_sec\": round(prompt_tps, 2) if prompt_tps else None,\n        \"eval_count\": eval_count,\n        \"eval_sec\": round(eval_duration / 1e9, 3) if eval_duration else None,\n        \"tokens_per_sec\": round(tps, 2) if tps else None,\n        \"response_chars\": len(response),\n        \"preview\": response[:160].replace(\"\\n\", \" \") + (\"...\" if len(response) > 160 else \"\"),\n    }\n\n\ndef run_benchmark(models: List[str], prompt: str, rounds: int, warmup: bool, think: str | None) -> None:\n    if warmup:\n        for model in models:\n            try:\n                call_ollama(model, \"Please answer in one sentence: test.\", think=think)\n            except Exception as e:\n                print(f\"[WARMUP FAIL] {model}: {e}\")\n\n    for model in models:\n        results = []\n        for i in range(rounds):\n            try:\n                result = call_ollama(model, prompt, think=think)\n                results.append(result)\n                print(json.dumps(result, ensure_ascii=False))\n            except Exception as e:\n                print(json.dumps({\"model\": model, \"round\": i + 1, \"error\": str(e)}, ensure_ascii=False))\n\n        speeds = [r[\"tokens_per_sec\"] for r in results if r.get(\"tokens_per_sec\") is not None]\n        if speeds:\n            summary = {\n                \"model\": model,\n                \"rounds\": len(speeds),\n                \"tokens_per_sec_avg\": round(statistics.mean(speeds), 2),\n                \"tokens_per_sec_min\": round(min(speeds), 2),\n                \"tokens_per_sec_max\": round(max(speeds), 2),\n            }\n            print(\"[SUMMARY] \" + json.dumps(summary, ensure_ascii=False))\n\n\nif __name__ == \"__main__\":\n    parser = argparse.ArgumentParser(description=\"Ollama API benchmark\")\n    parser.add_argument(\"models\", nargs=\"+\", help=\"Models to benchmark\")\n    parser.add_argument(\"--prompt\", default=\"Write a ~500-word explanation of practical use cases for local large language models.\")\n    parser.add_argument(\"--rounds\", type=int, default=1)\n    parser.add_argument(\"--warmup\", action=\"store_true\")\n    parser.add_argument(\"--think\", default=None, help=\"true / false / low / medium / high, if supported by model and Ollama\")\n    args = parser.parse_args()\n\n    run_benchmark(args.models, args.prompt, args.rounds, args.warmup, args.think)\n```\n\nUsage examples:\n\n```bash\npython3 ollama_benchmark.py gemma4:26b qwen3:14b --rounds 2 --warmup\npython3 ollama_benchmark.py mymodel-nothink mymodel --think false --rounds 2 --warmup\n```\n\n### 6.4 Quality Assessment Requires Real Tasks\n\nSpeed testing cannot replace quality testing. Before using a model in production workflows, test with real or anonymized materials:\n\n- Does it omit facts?\n- Does it fabricate information?\n- Can it follow fixed output formats?\n- Can it handle long texts?\n- Does it tend to over-infer?\n- Does it output thinking processes when it shouldn't?\n- Does it match the required task persona?\n\n---\n\n## 7. Creating Model Aliases\n\n### 7.1 Ollama Model Aliases\n\n```bash\nollama cp <source_model> <alias_name>\n```\n\nExample:\n\n```bash\nollama cp qwen3.6:35b-a3b-q4_k_m qq36\n```\n\nNotes:\n\n- `ollama cp` aliases typically share the underlying model storage.\n- Deleting an alias does not delete the source model.\n- Before deleting a source model, confirm whether its aliases are still needed.\n- Alias names should be short, stable, and reflect the model family or task purpose.\n\n### 7.2 Shell Command Aliases\n\nOnly suggest commands; do not automatically write to shell configuration.\n\n```bash\nalias qq36='ollama run qq36'\nalias g26='ollama run gemma4:26b'\n```\n\nIf the user wants to persist:\n\n```bash\necho \"alias qq36='ollama run qq36'\" >> ~/.zshrc\nsource ~/.zshrc\n```\n\nBefore writing, warn the user: this will modify the shell configuration file.\n\n---\n\n## 8. no-think / Thinking Control\n\n### 8.1 Key Principle\n\nA SYSTEM prompt can only ask the model to \"not output thinking process\" — this does not equal disabling thinking at the underlying level.\n\nIf Ollama and the model support the `think` parameter, prefer using the API to set it explicitly:\n\n```bash\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"Summarize this passage.\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\nIf the model supports thinking levels, you can also test:\n\n```json\n\"think\": \"low\"\n```\n\nor:\n\n```json\n\"think\": \"medium\"\n```\n\nWhether this takes effect depends on the model's actual behavior and Ollama API support.\n\n### 8.2 Modelfile Approach Is a Weak Constraint\n\nFor models that do not support the `think` parameter, a Modelfile can reduce explicit thinking but cannot guarantee thinking is truly disabled.\n\n```text\nFROM <model_name>:<tag>\nSYSTEM \"Output the final answer directly. Do not output thinking process, reasoning drafts, or hidden analysis.\"\nPARAMETER temperature 0.4\n```\n\nCreate:\n\n```bash\nollama create <model_name>-nothink -f Modelfile.<model_name>-nothink\n```\n\nRecommended file naming:\n\n```text\nModelfile.<source_model>.<purpose>.nothink\n```\n\nDo not overwrite existing Modelfiles.\n\n---\n\n## 9. Cleaning Up Old Models\n\n### 9.1 View Models and Disk Usage\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n### 9.2 Cleanup Criteria\n\nDo not mechanically \"keep only 3–5 models.\" Better criteria:\n\n- No clear purpose\n- No script references\n- Not an embedding / vision / coding / RAG specialty model\n- Not used in the last 30 days or confirmed no longer needed\n- Benchmark clearly behind peers\n- Already has a better alternative model\n- Can be re-downloaded or re-created after deletion\n\n### 9.3 Pre-Deletion Checklist\n\n```markdown\n## Pre-Deletion Check\n\n- Model to delete:\n- Current purpose:\n- In registry: Yes / No\n- Script references found: Yes / No\n- Is an alias: Yes / No\n- Alternative model available:\n- Deletion risk: Low / Medium / High\n- Recovery method: Re-pull / Re-create / Not recoverable\n- Recommendation: Delete / Defer / Keep\n```\n\nDelete command:\n\n```bash\nollama rm <model_name>\n```\n\nPost-deletion verification:\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\n---\n\n## 10. Quantization Level Selection Guide\n\n| Quantization | Size | Quality | Speed/Memory | Best For |\n|---|---:|---|---|---|\n| Q3_K_M | Small | Fair | Low pressure | Memory-constrained, rough tasks |\n| Q4_K_M | Medium | Good | Balanced | Daily default |\n| Q5_K_M | Larger | Better | Slightly higher | Quality-sensitive tasks |\n| Q8_0 | Large | Near-original | High pressure | Quality-first with ample memory |\n\nSelection principles:\n\n- General chat, summarization, light workflows: prefer Q4_K_M.\n- High-quality long-form, complex analysis: consider Q5_K_M or Q8_0.\n- Small memory devices: try Q3_K_M or smaller models first.\n- Different quantization levels must be compared with the same test set — do not judge by feel.\n\n---\n\n## 11. Common Issues\n\n### 11.1 Model Loading Failure Due to Old Ollama Version\n\nSymptoms:\n\n```text\nunable to load model\n```\n\nPossible causes:\n\n- Ollama version does not support the newer GGUF format\n- CLI and server versions mismatch\n- Server not restarted\n\nResolution:\n\n```bash\nollama --version\ncurl http://localhost:11434/api/version\n```\n\nThen upgrade or restart Ollama based on your installation method.\n\n### 11.2 Model Download Stuck\n\nResolution order:\n\n1. Confirm network is functional.\n2. Re-run `ollama pull` — Ollama typically supports resume.\n3. If the official source is stuck long-term, consider ModelScope or Hugging Face.\n4. After switching sources, re-validate model quality and format.\n\n### 11.3 Slow Model Loading or Insufficient Memory\n\nCheck:\n\n```bash\nollama list\ndu -sh ~/.ollama/models/\n```\n\nRecommendations:\n\n- Reduce model size or quantization level.\n- Reduce concurrency.\n- Avoid keeping multiple large models resident simultaneously.\n- Do not blindly increase `OLLAMA_CONTEXT_LENGTH`.\n- Only enable long context when the task actually requires it.\n\n### 11.4 Unstable Output Format\n\nPossible causes:\n\n- Model is not suited for structured output\n- Temperature too high\n- Prompt too loose\n- Context too long, diluting instructions\n\nResolution:\n\n- Lower temperature.\n- Add fixed templates.\n- Test format-following ability with short samples.\n- Replace the model if necessary.\n\n---\n\n## 12. Recommended Workflows\n\n### New Model Onboarding\n\n```text\nDiscover model → Check source → Determine purpose → Pull/Import → Record environment → Benchmark → Real-task testing → Add to registry → Decide whether to bind scripts\n```\n\n### Model Replacement\n\n```text\nTest candidate → Compare with current model → Check output quality → Small-scale replacement → Keep old model for observation → Update registry once stable → Then decide whether to delete old model\n```\n\n### Model Cleanup\n\n```text\nollama list → Cross-check registry → Scan script references → Assess alternatives → Output deletion suggestions → User confirms → Delete → Verify freed space → Update registry\n```\n\n---\n\n## 13. Quick Reference Commands\n\n```bash\n# Environment\nollama --version\ncurl http://localhost:11434/api/version\nollama list\ndu -sh ~/.ollama/models/\n\n# Pull model\nollama pull <model_name>:<tag>\n\n# ModelScope fallback\nollama pull modelscope.cn/<org>/<model_name>:<tag>\n\n# Create Ollama alias\nollama cp <source_model> <alias_name>\n\n# Scan script references\nrg \"ollama run|MODEL_NAME|model_name|\\\"model\\\"|ollama\\.chat|ollama\\.generate\" .\n\n# Delete model (confirm before executing)\nollama rm <model_name>\n\n# API no-think test\ncurl http://localhost:11434/api/generate -d '{\n  \"model\": \"<model>\",\n  \"prompt\": \"Please answer in one sentence: test.\",\n  \"think\": false,\n  \"stream\": false\n}'\n```\n\n---\n\n## 14. Output Requirements\n\nWhen a user requests \"help me manage Ollama models\", prioritize outputting:\n\n1. Current assessment\n2. Safety risks\n3. Recommended read-only checks\n4. Whether user confirmation is needed\n5. Next commands\n\nWhen a user requests \"clean up models\", output:\n\n1. Model tiering\n2. Non-deletable models\n3. Models under observation\n4. Deletion candidates\n5. Pre-deletion reference scan command\n6. Deletion confirmation prompt\n\nWhen a user requests \"test models\", output:\n\n1. Testing objectives\n2. Test scenarios\n3. Benchmark commands\n4. Quality assessment dimensions\n5. Whether to recommend for production workflow","readmeExcerpt":"Skill: Ollama Model Pilot Owner: patmenciu Summary: Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-th... Tags: latest:1.5.0 Version history: v1.5.0 | 2026-06-08T14:26:51.217Z | user **Modelpilot v1.5.0 Changelog** - Major refactor: SKILL.md now emphasizes a strict local-only safety boundary, two-round replaceme","codeSnippets":[],"executableExamples":[{"language":"markdown","snippet":"## ModelPilot Result\n\n### Scope\n-\n\n### Models Tested\n-\n\n### Test Rounds\n-\n\n### Key Findings\n-\n\n### No-Think Check\n-\n\n### Replacement Decision\n-\n\n### Risks and Limits\n-\n\n### Rollback Advice\n-"},{"language":"text","snippet":"modelpilot/\n  SKILL.md\n  README.md\n  scripts/\n  examples/\n  tests/\n  outputs/"},{"language":"bash","snippet":"python scripts/modelpilot_benchmark.py \\\n  --models llama3.2:latest qwen3:latest \\\n  --prompts examples/prompts.example.json \\\n  --rounds 2 \\\n  --output outputs/benchmark_results.json"},{"language":"bash","snippet":"python scripts/modelpilot_report.py \\\n  --input outputs/benchmark_results.json \\\n  --output outputs/benchmark_report.md"},{"language":"bash","snippet":"curl http://localhost:11434/api/version"},{"language":"bash","snippet":"curl http://localhost:11434/api/tags"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: modelpilot\ndescription: Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-think verification, and local-only safety boundaries. It applies to local LLM evaluation, model replacement decisions, benchmark reports, installed-model audits, and Ollama workflow hygiene. Do not use it for cloud model APIs, downloading models, installing dependencies, or sending local data outside the machine.\n---\n\n# ModelPilot\n\nModelPilot is a local-only protocol for testing, comparing, promoting, replacing,\nand cleaning up Ollama models. It is designed for real work decisions, not leaderboard\nclaims.\n\n## Safety Boundary\n\nAlways keep the workflow local unless the user explicitly authorizes otherwise.\n\n- Do not call cloud model APIs.\n- Do not upload files, prompts, logs, paths, configs, or benchmark outputs.\n- Do not download, pull, install, upgrade, or delete models without explicit user approval.\n- Do not use real private documents as benchmark samples unless the user explicitly names the file for this task.\n- Use fictional examples for tests, documentation, and demos.\n- Treat model cleanup as a workflow dependency audit, not a disk-space optimization task.\n\n## Trigger Conditions\n\nUse this skill when the user asks to:\n\n- test an Ollama model\n- compare local models\n- decide whether a new model can replace an existing model\n- verify no-think behavior\n- build a local model benchmark report\n- audit installed models before cleanup\n- choose local models for coding, writing, RAG, automation, or structured output\n\n## Test Levels\n\nClassify the task before running anything.\n\n1. Smoke Test\n   Confirm the model is installed, runnable, and responsive.\n\n2. Speed Benchmark\n   Measure startup time, generation time, output length, and failure rate.\n\n3. Real-Task Benchmark\n   Use task-like prompts that match the user's actual workflow. Prefer fixed prompt\n   sets so results are comparable across models.\n\n4. Promotion Test\n   Decide whether a model can replace an existing workflow model. A promotion test\n   requires two independent benchmark rounds.\n\n## Two-Round Replacement Rule\n\nDo not recommend replacing a working model after a single run.\n\n- Round 1 checks: runnable, speed, output format, obvious quality failures, no-think leakage.\n- Round 2 checks: same prompt set, same model, repeatability, quality consistency, failure modes.\n- A model is only replacement-ready when both rounds pass the required tasks.\n- Keep the previous model and configuration available for rollback.\n- If structured output, no-think behavior, or long-context handling is unstable, do not use the model in automation.\n\n## Fixed Prompt Set\n\nPrefer a stable prompt file with fictional data. Include at least:\n\n- short Chinese or English Q&A\n- long-document summary\n- structured JSON or Markdown output\n- real-role workflow simulation\n- no-think verification prompt\n\nThe benchmark prompt set should be reused ac"},{"path":"README.md","content":"# ModelPilot\n\nModelPilot is a local-only skill for testing, comparing, promoting, replacing,\nand cleaning up Ollama models.\n\nIt is built around a simple rule: a model should not replace an existing workflow\nmodel after one good run. Run the same fixed prompt set twice, review both rounds,\nthen decide.\n\n## What It Helps With\n\n- Compare local Ollama models on real tasks\n- Verify whether a `nothink` model actually suppresses thinking traces\n- Decide whether a candidate model can replace a current model\n- Produce compact benchmark reports\n- Audit models before cleanup without deleting anything automatically\n\n## Directory Layout\n\n```text\nmodelpilot/\n  SKILL.md\n  README.md\n  scripts/\n  examples/\n  tests/\n  outputs/\n```\n\n## Safety Defaults\n\n- Local Ollama only\n- No cloud model APIs\n- No uploads\n- No model downloads\n- No dependency installation\n- No automatic model deletion\n- Fictional examples only\n\n## Basic Usage\n\nPrepare a fictional prompt set:\n\n```bash\npython scripts/modelpilot_benchmark.py \\\n  --models llama3.2:latest qwen3:latest \\\n  --prompts examples/prompts.example.json \\\n  --rounds 2 \\\n  --output outputs/benchmark_results.json\n```\n\nCreate a Markdown report:\n\n```bash\npython scripts/modelpilot_report.py \\\n  --input outputs/benchmark_results.json \\\n  --output outputs/benchmark_report.md\n```\n\nThe scripts use Python standard library only. The benchmark script calls the local\n`ollama` command and expects the models to already be installed.\n\n## Replacement Rule\n\nA model can be considered replacement-ready only after two independent rounds using\nthe same fixed prompt set.\n\nRound 1 checks:\n\n- the model runs\n- response speed is acceptable\n- output format is stable\n- no-think behavior does not leak reasoning text\n\nRound 2 checks:\n\n- the same tasks still pass\n- failure modes do not repeat\n- quality is consistent enough for the target workflow\n\nIf either round fails on structured output, no-think behavior, or the user's core\ntask, keep the existing model.\n\n## Example Decision Labels\n\n- `replace_ready`: both rounds pass and manual review confirms quality\n- `observe`: usable, but has minor instability or incomplete evidence\n- `candidate_only`: only one round is complete\n- `not_recommended`: repeated failures, output pollution, or unsafe automation fit\n\n## Limits\n\nModelPilot does not prove general intelligence or leaderboard quality. It helps make\nlocal workflow decisions based on fixed tasks, repeatability, and clean output.\n\nThe report script can detect mechanical issues, but semantic quality still needs a\nhuman review."},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7fj0qnbect2fmp59jj0vx34587hvns\",\n  \"slug\": \"modelpilot\",\n  \"version\": \"1.5.0\",\n  \"publishedAt\": 1780928811217\n}"},{"path":"outputs/example_report.md","content":"# ModelPilot Benchmark Report\n\n- Version: 1.5.0\n- Generated at: 2026-06-08T00:00:00+00:00\n- Local only: true\n- Rounds requested: 2\n- Prompt count: 5\n\n## Summary\n\n| Model | Decision | Reason | Success | Format | Think leak | Avg seconds |\n| --- | --- | --- | ---: | ---: | ---: | ---: |\n| example-candidate-a:latest | replace_ready | Two rounds passed mechanical checks. Human semantic review is still required. | 10/10 | 10/10 | 0 | 3.42 |\n| example-candidate-b:nothink | not_recommended | 1 outputs show possible thinking leakage. | 10/10 | 10/10 | 1 | 3.10 |\n\n## Replacement Decision\n\n- `example-candidate-a:latest` can be considered for replacement after manual semantic review.\n- `example-candidate-b:nothink` should not be promoted into automated workflows until no-think output is clean.\n- Keep the previous model and config available for rollback.\n\n## Risks and Limits\n\n- This example uses fictional data.\n- Mechanical checks do not prove semantic quality.\n- Do not delete old models based only on a benchmark report."},{"path":"skill-card.md","content":"## Description:\n\nModelPilot helps agents test, compare, promote, replace, and clean up local Ollama models using repeatable two-round benchmarks, no-think checks, and local-only safety boundaries.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[patmenciu](https://clawhub.ai/user/patmenciu)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and engineers use this skill to evaluate installed local Ollama models on repeatable task prompts, no-think behavior, structured output stability, and replacement readiness. It supports local workflow decisions without cloud APIs, uploads, model downloads, or automatic model deletion.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Benchmark reports can overstate replacement readiness because mechanical checks do not prove semantic quality.\n\nMitigation: Require human semantic review and two clean benchmark rounds before replacing a workflow model; keep the previous model and configuration available for rollback.\n\nRisk: Prompts, logs, or benchmark outputs may contain sensitive local data if real files are used.\n\nMitigation: Use fictional prompt sets by default, only use user-named files for the task, and keep benchmark outputs local.\n\nRisk: No-think model names or instructions may not guarantee clean output.\n\nMitigation: Check outputs for thinking traces and avoid promoting a model into automation when leakage appears.\n\nRisk: Untrusted or edited benchmark JSON can produce misleading reports.\n\nMitigation: Review benchmark inputs and reports manually, and rerun benchmarks from trusted model and prompt configurations when results affect replacement decisions.\n\n## Reference(s):\n\n- [ClawHub Skill Page](https://clawhub.ai/patmenciu/skills/modelpilot)\n- [README](README.md)\n- [Example Benchmark Report](outputs/example_report.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with optional shell commands, JSON benchmark results, and Markdown reports]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Local-only Ollama workflow; benchmark scripts use installed models and fictional or explicitly selected prompt files.]\n\n## Skill Version(s):\n\n1.5.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1940,"uniquenessScore":38,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T04:31:17.124Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T04:31:17.124Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T07:40:11.858Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}