agentCLAWHUBUnverified

llama-params-optimizer

Complete methodology for local LLM performance optimization. Core principle: maximize context while fully covering GPU memory — find the sweet spot where GPU... Skill: llama-params-optimizer Owner: hoperealize Summary: Complete methodology for local LLM performance optimization. Core principle: maximize context while fully covering GPU memory — find the sweet spot where GPU... Tags: latest:3.1.0, llama.cpp:2.0.0, local-llm:2.0.0, optimization:2.0.0, performance:2.0.0 Version history: v3.1.0 | 2026-04-28T11:07:43.048Z | user 新增 Qwen3.6-27B 实战案例;修正甜点公式为起点参考;batch-size 因模型而异;推理

OpenClaw

Rank

62

Safety

84

Downloads

1.1k

Updated

Oct 11, 2026

Version

3.1.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.1K downloadsadoption · observed Oct 11, 2026
Latest release
3.1.0release · observed Apr 28, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17ad6sxtk65d4y20pyavaxmys83mb6w:llama-params-optimizer
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-hoperealize-llama-params-optimizer/snapshot"

Documentation

CLAWHUB

145,855 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: llama-params-optimizer
version: 3.1.0
description: >
  Complete methodology for local LLM performance optimization.
  Core principle: maximize context while fully covering GPU memory — find the sweet spot where GPU runs at full speed.
  Step-by-step 4-phase 10-step control variable testing process.
  Works for ALL llama.cpp / llama-server models on ANY hardware.
  Cases: Qwen3.5-MoE, Qwen3.6-35B, Qwen3.6-27B (2026-04-28).
author: fenglai
keywords: [llama.cpp, performance optimization, local llm, llama-server, quantization, long context, control variable testing, speed optimization, reasoning models, OpenClaw config]
tags: [llm, performance, optimization, local-first, chinese-support]
---

# llama.cpp 启动参数优化技能 / llama.cpp Parameter Optimization Guide

**中文** | **English**  
标准化的 LLM 本地部署启动参数优化评估流程,通过严格的控制变量测试,找到最佳的性能/质量平衡点。  
_A standardized methodology for optimizing local LLM deployment parameters, using rigorous control variable testing to find the optimal performance/quality balance._

**⚠️ 安全声明 / Safety Notice**
- 本技能仅用于 **本地部署** 参数优化,所有测试应在隔离环境中进行。
- **llama-server 和 llama.cpp 二进制文件**应仅从官方 GitHub 仓库 (https://github.com/ggerganov/llama.cpp) 获取,避免使用第三方来源。
- **网络绑定安全**:本地开发/测试时建议绑定 `127.0.0.1`(localhost),生产环境必须通过反向代理 + HTTPS 暴露服务。
- 参数调优可能触发 OOM 崩溃,**不确定最优参数时,建议使用云端模型(如 OpenAI/Gemini)进行推理验证**,本地只用于性能调优。
- **不要在生产环境中使用本技能提供的示例命令直接暴露服务**,需根据实际安全需求调整。

## 🎯 核心卖点 / Key Features
- ✅ **GPU 内存覆盖原则** / _GPU Memory Coverage Principle_:在完全使用专用 GPU 内存的前提下,找到最大上下文值 | _Find max context while fully covering GPU memory_
- ✅ **四阶段十步法控制变量测试** / _4-phase 10-step process_:完整的性能/质量评估 | _Comprehensive performance and quality evaluation_
- ✅ **大量反常识踩坑经验** / _Battle-tested counterintuitive findings_:避免踩同样的坑 | _Avoid common pitfalls_
- ✅ **通用方法论** / _Universal methodology_:提供系统化测试框架,具体参数需结合实际硬件验证 | _Provides a systematic testing framework; specific parameters must be validated against actual hardware._
- ✅ **实战验证** / _Real-world proven_:35B+4060Ti 案例:~30 → ~90 token/s,提升 2-3 倍;27B 案例:~23.6 tok/s,验证理论公式不可靠 | _Multiple cases: MoE 262% boost, Dense 2-3x, 27B theoretical formula fails_

## 适用场景 / When to Use

**中文** | **English**  
新模型首次部署,需要找到最佳启动参数  
_First-time deployment of a new model, finding optimal launch parameters_  
新硬件环境下的性能调优  
_Performance tuning on new hardware_  
llama.cpp / llama-server 启动参数优化  
_llama.cpp / llama-server launch parameter optimization_  
验证量化损失、长上下文能力等核心特性  
_Verify quantization loss, long context capabilities, and other core features_

---

## 完整评估流程 / Complete Methodology
_**4 phases, 10 steps**_

---

### 📊 第一阶段:基准建立 / Phase 1: Establish Baseline

#### 步骤 1:建立初始基准 / Step 1: Run at default parameters
**中文** | **English**  
在默认参数下运行,记录基础性能数据:  
_Run with default parameters and record baseline performance:_  
```
✅ 记录项 / Metrics to record:
- 生成速度 / Generation speed (tokens/s)
- Prompt 处理速度 / Prompt processing speed (tokens/s)
- 显存占用峰值 / Peak VRAM usage (GB)
- 首字延迟 / Time to first token (ms)
```

#### 步骤 2:枚举所有待测试参数 / Step 2: Li

_meta.json

{
  "ownerId": "kn7ers0mgs9t4jm7cy6s7fpktd82kzr0",
  "slug": "llama-params-optimizer",
  "version": "3.1.0",
  "publishedAt": 1777374463048
}

CHANGELOG.md

# CHANGELOG

## v3.1.0 (2026-04-28)

### ✨ 新增
- 新增 Qwen3.6-27B Dense + RTX 4060Ti 16GB 实战案例
- 新增推理模型(reasoning models)思考开销说明:Qwen3.6 内置推理链,TTFT 40-50 秒
- 新增 OpenClaw 配置对齐指南:contextWindow/maxTokens 必须与 --ctx-size 匹配

### 🐛 修正
- 甜点公式从"唯一真理"降级为"起点参考",强调理论计算不可靠,实测才是王道(27B 上公式算 110K OOM,实际甜点 96K)
- batch-size 最优值因模型大小而异:35B Dense 上 2048 最优,27B Dense 上 512 反而快 6.6%
- 核心原则从 7 条扩展至 10 条

---

## v1.2.0 (2026-04-26)

### ✨ 国际化
- 完整中英文双语版本(Bilingual),全球用户可用
- Frontmatter description 改为英文,便于搜索引擎索引
- 所有章节标题、核心概念、步骤说明均提供中英文对照
- 保留完整的中文详细步骤说明

---

## v1.1.2 (2026-04-26)

### 🐛 修正
- 将实战案例拆分为两个:MoE 模型(262%提升)和 Dense 模型(77%提升),避免参数混淆
- 每个案例单独列出最佳参数,明确区分不同模型类型的差异
- 补充 120K 甜点阈值的完整对比数据表格

---

## v1.1.1 (2026-04-26)

### 🐛 修正
- 更新实战案例的最终启动命令,从 64K 修正为实测最优的 120K 甜点阈值
- 修正 `--parallel` 参数的推荐说明:Dense 35B 模型上 parallel=2 反而慢 40%,建议保持 1
- 补充 Linux/WSL2 的 CPU 亲和性绑定启动命令

---

## v1.1.0 (2026-04-26)

### ✨ 重大更新
- 新增「上下文甜点阈值」优先测试章节,这是目前性价比最高的优化方法
  - 详细解释了断崖式性能下降的原理(显存Bank对齐、FlashAttention块大小、大页内存)
  - 标准化的4步测试方法,任何模型/硬件都可以复用
  - 附带 Qwen3.6-35B + RTX 4060Ti 的完整测试案例
  - 上下文仅少6%,速度提升75%,零质量损失

### 📝 更新反常识发现列表
- ❌ `--parallel 2` 不一定好:在 4060Ti + 35B 组合上,单请求速度反而下降 40%
- ✅ `--flash-attn on` 对长 Prompt 影响巨大:3-5 倍提升
- ❌ q4_K KV 缓存可能有兼容性问题:加载速度极慢,优先用 q8_0

### 🆕 新增内容
- CPU 亲和性绑定章节:+1-5% 免费提升,Linux/WSL2 适用

---

## v1.0.0 (2026-04-26)

### ✨ 初始版本
- 完整的四阶段十步法控制变量测试流程
- Qwen3.5-MoE 35B 实战案例,性能提升 262%
- 所有关键参数的最佳实践和测试方法
- 长上下文召回验证、量化损失验证等质量测试方法
- 多维度综合评分方法

example-qwen3.5-moe-4060ti.md

# 实战案例:Qwen3.5-MoE 35B + RTX 4060Ti 16GB

## 测试环境
- 模型:Qwen3.6-35B-A3B-APEX-I-Mini.gguf(Q3_K_M,13.3GB)
- 显卡:NVIDIA RTX 4060Ti 16GB
- CPU:i5-14600KF
- 软件:llama.cpp b8925

---

## 优化成果
**初始速度:23.4 tokens/s → 最终速度:84.8 tokens/s,提升 262%!**

---

## 完整测试数据

### 1. 线程数测试
| 线程数 | 生成速度 | Prompt 速度 | 生成速度变化 | 推荐 |
|--------|---------|------------|------------|------|
| 8 | 84.8 | 80.0 | 基准 | 🏆 最佳 |
| 12 | 83.1 | 70.2 | -2.0% | |
| 16 | 83.5 | 75.0 | -1.5% | |

**结论:线程不是越多越好,8 线程最佳。**

---

### 2. Batch Size 测试
| batch size | 生成速度 | Prompt 速度 | 生成速度变化 | 推荐 |
|------------|---------|------------|------------|------|
| 512(默认) | 51.2 | 35.0 | 基准 | ❌ 太慢! |
| 1024 | 54.3 | 35.5 | +6.1% | |
| 2048 | 84.8 | 80.0 | +67.7% | 🏆 最佳 |
| 4096 | 85.3 | 85.0 | +67.9% | |

**重大发现:默认 batch size 只有 512,改成 2048 直接快 67.7%!这是本次最大的性能提升点!**

---

### 3. Flash Attention + KV 量化测试
| 配置 | 生成速度 | Prompt 速度 | 生成变化 | Prompt 变化 | 推荐 |
|------|---------|------------|---------|------------|------|
| FA off + 不量化 | 85.9 | 35.0 | 基准 | 基准 | |
| FA on + KV q8_0 | 84.8 | 80.0 | -1.3% | +128.6% | 🏆 最佳 |
| FA on + KV q4_0 | 84.8 | 59.4 | -0.0% | +69.7% | |

**反常识发现:**
1. KV 量化不是损失!8bit KV 反而让 Prompt 处理快了 128%!
2. Flash Attention 对 MoE 模型:生成只慢 1.3%,但 Prompt 快 128%,整体收益巨大!
3. q8_0 是最佳平衡点,q4_0 的 Prompt 速度下降明显。

---

### 4. 上下文窗口测试
| ctx-size | 生成速度 | Prompt 速度 | 生成变化 | 推荐 |
|----------|---------|------------|---------|------|
| 65536(64K) | 84.8 | 80.0 | 基准 | 🏆 速度优先 |
| 131072(128K) | 24.0 | 45.4 | -71.7% | |
| 262144(256K) | 22.9 | 32.7 | -73.0% | 📚 上下文优先 |

**结论:上下文翻倍,速度几乎减半。根据使用场景选择。**

---

### 5. 并行会话数测试
| --parallel | 生成速度 | Prompt 速度 | 生成变化 | Prompt 变化 | 推荐 |
|------------|---------|------------|---------|------------|------|
| 1 | 86.6 | 64.4 | +2.1% | -19.5% | |
| 2(默认) | 84.8 | 80.0 | 基准 | 基准 | 🏆 最佳 |
| 4 | 84.8 | 46.5 | +0.0% | -41.9% | ❌ |

**结论:--parallel 2 是最佳平衡点,Prompt 速度最快,还支持 2 并发。**

---

### 6. 微批量大小测试
| ubatch-size | 生成速度 | 变化 | 推荐 |
|--------------|---------|------|------|
| 256 | 70.9 | -17.5% | ❌ |
| 512(默认) | 84.8 | 基准 | 🏆 别动! |
| 1024 | 34.2 | -60.2% | ❌ |

**重大警告:千万别改 ubatch-size!默认的 512 就是最佳的,改了反而慢 17-60%!**

---

## 质量验证结果

| 测试项 | 结果 |
|--------|------|
| KV q8_0 量化质量对比 | ✅ 无任何可感知的质量损失 |
| 短距离回忆(1000 token) | ✅ 100% 正确 |
| 中距离回忆(40000 token) | ✅ 100% 正确 |
| 简单数学/推理 | ✅ 100% 正确 |
| 简单代码生成 | ✅ 正常 |

---

## 综合评分

| 维度 | 权重 | 得分 |
|------|------|------|
| 性能 | 50% | 9.5/10 |
| 质量 | 40% | 9.4/10 |
| 稳定性 | 10% | 10.0/10 |
| **总分** | 100% | **9.5/10** 🏆 |

---

## 最终最佳配置

### ⚡ 方案一:速度优先(84.8 tokens/s + 80 tokens/s Prompt)
```cmd
llama-server.exe -m "Qwen3.6-35B-A3B-APEX-I-Mini.gguf" --n-gpu-layers 9999 --ctx-size 65536 --port 8080 --host 127.0.0.1 --threads 8 --mlock --parallel 2 --kv-unified --flash-attn on -b 2048 --cache-type-k q8_0 --cache-type-v q8_0
```

### 📚 方案二:上下文优先(256K 超大窗口)
```cmd
llama-server.exe -m "Qwen3.6-35B-A3B-APEX-I-Mini.gguf" --n-gpu-layers 9999 --ctx-size 262144 --port 8080 --host 127.0.0.1 --threads 8 --mlock --parallel 2 --kv-unified --flas

skill-card.md

## Description:

Complete methodology for local LLM performance optimization using a four-phase control-variable process to tune llama.cpp and llama-server parameters across hardware.

This skill is ready for commercial/non-commercial use.

## Publisher:

[hoperealize](https://clawhub.ai/user/hoperealize)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and engineers use this skill to benchmark local llama.cpp and llama-server deployments, compare launch parameters, and produce recommended configurations for performance, context length, and stability.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Example llama-server commands could expose a local model service if copied into an externally reachable deployment.

Mitigation: Keep llama-server bound to localhost unless a reverse proxy, HTTPS, and deployment-specific access controls are configured.

Risk: Parameter tuning can trigger out-of-memory failures or mismatched context settings on local hardware.

Mitigation: Validate settings against the target GPU and align OpenClaw contextWindow and maxTokens with the selected llama-server context size.

Risk: Cloud model checks used during tuning may receive sensitive prompts if operators reuse private test data.

Mitigation: Use non-sensitive prompts for cloud validation and keep private or regulated content out of external model calls.

## Reference(s):

- [llama.cpp](https://github.com/ggerganov/llama.cpp)
- [ClawHub skill page](https://clawhub.ai/hoperealize/skills/llama-params-optimizer)

## Skill Output:

**Output Type(s):** [Text, Markdown, Shell commands, Configuration, Guidance]

**Output Format:** [Markdown with benchmark tables, command examples, and configuration snippets]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Produces hardware-specific recommendations that should be validated on the target machine.]

## Skill Version(s):

3.1.0 (source: frontmatter, changelog, server release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 2d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/hoperealize/skills/llama-params-optimizer",
      "sourceUrl": "https://clawhub.ai/hoperealize/skills/llama-params-optimizer",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T11:35:50.761Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-hoperealize-llama-params-optimizer/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-hoperealize-llama-params-optimizer/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T11:35:50.761Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.1K downloads",
      "href": "https://clawhub.ai/hoperealize/llama-params-optimizer",
      "sourceUrl": "https://clawhub.ai/hoperealize/llama-params-optimizer",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T11:35:50.761Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "3.1.0",
      "href": "https://clawhub.ai/hoperealize/llama-params-optimizer",
      "sourceUrl": "https://clawhub.ai/hoperealize/llama-params-optimizer",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-28T11:07:43.048Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-hoperealize-llama-params-optimizer/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-hoperealize-llama-params-optimizer/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 3.1.0",
      "description": "新增 Qwen3.6-27B 实战案例;修正甜点公式为起点参考;batch-size 因模型而异;推理模型 TTFT 开销说明;OpenClaw 配置对齐指南",
      "href": "https://clawhub.ai/hoperealize/llama-params-optimizer",
      "sourceUrl": "https://clawhub.ai/hoperealize/llama-params-optimizer",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-28T11:07:43.048Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to llama-params-optimizer and adjacent AI workflows.