Ollama Model Pilot
Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-th...
Rank
62
Safety
84
Downloads
1.2k
Updated
Oct 11, 2026
Version
1.5.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.2K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.5.0release · observed Jun 8, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17b7ke0py60etvzaczrwdabh587gytb:modelpilot- Install using `clawhub skill install s17b7ke0py60etvzaczrwdabh587gytb:modelpilot` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/patmenciu/modelpilot before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/snapshot"
Documentation
CLAWHUB
123,436 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: modelpilot description: Use this skill when the user wants to test, compare, promote, replace, or clean up local Ollama models with a repeatable two-round real-task benchmark, no-think verification, and local-only safety boundaries. It applies to local LLM evaluation, model replacement decisions, benchmark reports, installed-model audits, and Ollama workflow hygiene. Do not use it for cloud model APIs, downloading models, installing dependencies, or sending local data outside the machine. --- # ModelPilot ModelPilot is a local-only protocol for testing, comparing, promoting, replacing, and cleaning up Ollama models. It is designed for real work decisions, not leaderboard claims. ## Safety Boundary Always keep the workflow local unless the user explicitly authorizes otherwise. - Do not call cloud model APIs. - Do not upload files, prompts, logs, paths, configs, or benchmark outputs. - Do not download, pull, install, upgrade, or delete models without explicit user approval. - Do not use real private documents as benchmark samples unless the user explicitly names the file for this task. - Use fictional examples for tests, documentation, and demos. - Treat model cleanup as a workflow dependency audit, not a disk-space optimization task. ## Trigger Conditions Use this skill when the user asks to: - test an Ollama model - compare local models - decide whether a new model can replace an existing model - verify no-think behavior - build a local model benchmark report - audit installed models before cleanup - choose local models for coding, writing, RAG, automation, or structured output ## Test Levels Classify the task before running anything. 1. Smoke Test Confirm the model is installed, runnable, and responsive. 2. Speed Benchmark Measure startup time, generation time, output length, and failure rate. 3. Real-Task Benchmark Use task-like prompts that match the user's actual workflow. Prefer fixed prompt sets so results are comparable across models. 4. Promotion Test Decide whether a model can replace an existing workflow model. A promotion test requires two independent benchmark rounds. ## Two-Round Replacement Rule Do not recommend replacing a working model after a single run. - Round 1 checks: runnable, speed, output format, obvious quality failures, no-think leakage. - Round 2 checks: same prompt set, same model, repeatability, quality consistency, failure modes. - A model is only replacement-ready when both rounds pass the required tasks. - Keep the previous model and configuration available for rollback. - If structured output, no-think behavior, or long-context handling is unstable, do not use the model in automation. ## Fixed Prompt Set Prefer a stable prompt file with fictional data. Include at least: - short Chinese or English Q&A - long-document summary - structured JSON or Markdown output - real-role workflow simulation - no-think verification prompt The benchmark prompt set should be reused ac
README.md
# ModelPilot ModelPilot is a local-only skill for testing, comparing, promoting, replacing, and cleaning up Ollama models. It is built around a simple rule: a model should not replace an existing workflow model after one good run. Run the same fixed prompt set twice, review both rounds, then decide. ## What It Helps With - Compare local Ollama models on real tasks - Verify whether a `nothink` model actually suppresses thinking traces - Decide whether a candidate model can replace a current model - Produce compact benchmark reports - Audit models before cleanup without deleting anything automatically ## Directory Layout ```text modelpilot/ SKILL.md README.md scripts/ examples/ tests/ outputs/ ``` ## Safety Defaults - Local Ollama only - No cloud model APIs - No uploads - No model downloads - No dependency installation - No automatic model deletion - Fictional examples only ## Basic Usage Prepare a fictional prompt set: ```bash python scripts/modelpilot_benchmark.py \ --models llama3.2:latest qwen3:latest \ --prompts examples/prompts.example.json \ --rounds 2 \ --output outputs/benchmark_results.json ``` Create a Markdown report: ```bash python scripts/modelpilot_report.py \ --input outputs/benchmark_results.json \ --output outputs/benchmark_report.md ``` The scripts use Python standard library only. The benchmark script calls the local `ollama` command and expects the models to already be installed. ## Replacement Rule A model can be considered replacement-ready only after two independent rounds using the same fixed prompt set. Round 1 checks: - the model runs - response speed is acceptable - output format is stable - no-think behavior does not leak reasoning text Round 2 checks: - the same tasks still pass - failure modes do not repeat - quality is consistent enough for the target workflow If either round fails on structured output, no-think behavior, or the user's core task, keep the existing model. ## Example Decision Labels - `replace_ready`: both rounds pass and manual review confirms quality - `observe`: usable, but has minor instability or incomplete evidence - `candidate_only`: only one round is complete - `not_recommended`: repeated failures, output pollution, or unsafe automation fit ## Limits ModelPilot does not prove general intelligence or leaderboard quality. It helps make local workflow decisions based on fixed tasks, repeatability, and clean output. The report script can detect mechanical issues, but semantic quality still needs a human review.
_meta.json
{
"ownerId": "kn7fj0qnbect2fmp59jj0vx34587hvns",
"slug": "modelpilot",
"version": "1.5.0",
"publishedAt": 1780928811217
}outputs/example_report.md
# ModelPilot Benchmark Report - Version: 1.5.0 - Generated at: 2026-06-08T00:00:00+00:00 - Local only: true - Rounds requested: 2 - Prompt count: 5 ## Summary | Model | Decision | Reason | Success | Format | Think leak | Avg seconds | | --- | --- | --- | ---: | ---: | ---: | ---: | | example-candidate-a:latest | replace_ready | Two rounds passed mechanical checks. Human semantic review is still required. | 10/10 | 10/10 | 0 | 3.42 | | example-candidate-b:nothink | not_recommended | 1 outputs show possible thinking leakage. | 10/10 | 10/10 | 1 | 3.10 | ## Replacement Decision - `example-candidate-a:latest` can be considered for replacement after manual semantic review. - `example-candidate-b:nothink` should not be promoted into automated workflows until no-think output is clean. - Keep the previous model and config available for rollback. ## Risks and Limits - This example uses fictional data. - Mechanical checks do not prove semantic quality. - Do not delete old models based only on a benchmark report.
skill-card.md
## Description: ModelPilot helps agents test, compare, promote, replace, and clean up local Ollama models using repeatable two-round benchmarks, no-think checks, and local-only safety boundaries. This skill is ready for commercial/non-commercial use. ## Publisher: [patmenciu](https://clawhub.ai/user/patmenciu) ### License/Terms of Use: MIT-0 ## Use Case: Developers and engineers use this skill to evaluate installed local Ollama models on repeatable task prompts, no-think behavior, structured output stability, and replacement readiness. It supports local workflow decisions without cloud APIs, uploads, model downloads, or automatic model deletion. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Benchmark reports can overstate replacement readiness because mechanical checks do not prove semantic quality. Mitigation: Require human semantic review and two clean benchmark rounds before replacing a workflow model; keep the previous model and configuration available for rollback. Risk: Prompts, logs, or benchmark outputs may contain sensitive local data if real files are used. Mitigation: Use fictional prompt sets by default, only use user-named files for the task, and keep benchmark outputs local. Risk: No-think model names or instructions may not guarantee clean output. Mitigation: Check outputs for thinking traces and avoid promoting a model into automation when leakage appears. Risk: Untrusted or edited benchmark JSON can produce misleading reports. Mitigation: Review benchmark inputs and reports manually, and rerun benchmarks from trusted model and prompt configurations when results affect replacement decisions. ## Reference(s): - [ClawHub Skill Page](https://clawhub.ai/patmenciu/skills/modelpilot) - [README](README.md) - [Example Benchmark Report](outputs/example_report.md) ## Skill Output: **Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with optional shell commands, JSON benchmark results, and Markdown reports] **Output Parameters:** [1D] **Other Properties Related to Output:** [Local-only Ollama workflow; benchmark scripts use installed models and fictional or explicitly selected prompt files.] ## Skill Version(s): 1.5.0 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/patmenciu/skills/modelpilot",
"sourceUrl": "https://clawhub.ai/patmenciu/skills/modelpilot",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T04:31:17.124Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T04:31:17.124Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.2K downloads",
"href": "https://clawhub.ai/patmenciu/modelpilot",
"sourceUrl": "https://clawhub.ai/patmenciu/modelpilot",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T04:31:17.124Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.5.0",
"href": "https://clawhub.ai/patmenciu/modelpilot",
"sourceUrl": "https://clawhub.ai/patmenciu/modelpilot",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-06-08T14:26:51.217Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-patmenciu-modelpilot/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.5.0",
"description": "**Modelpilot v1.5.0 Changelog** - Major refactor: SKILL.md now emphasizes a strict local-only safety boundary, two-round replacement rule, and repeatable benchmark protocol for Ollama models. - Added: Benchmark example files (`examples/models.example.json`, `examples/prompts.example.json`, `outputs/example_report.md`) and local scripts for running/reporting benchmarks (`scripts/modelpilot_benchmark.py`, `scripts/modelpilot_report.py`). - Added: Unit test for the new reporting workflow (`tests/test_modelpilot_report.py`). - Removed: Skill metadata files (`CHANGELOG.md`, `skill-card.md`) now replaced by clearer in-skill documentation. - New: Required response format for benchmark and decision results to ensure consistent reporting. - Now explicitly prohibits cloud API calls, model downloading, and any non-local actions without user approval.",
"href": "https://clawhub.ai/patmenciu/modelpilot",
"sourceUrl": "https://clawhub.ai/patmenciu/modelpilot",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-06-08T14:26:51.217Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
