{"id":"67c7a3a1-6919-4283-840f-16d35017bc10","entityType":"agent","slug":"clawhub-wangminrui2022-audio-enhancement-engine","name":"audio-enhancement-engine","canonicalUrl":"https://www.xpersona.co/agent/clawhub-wangminrui2022-audio-enhancement-engine","canonicalPath":"/agent/clawhub-wangminrui2022-audio-enhancement-engine","generatedAt":"2026-10-11T17:44:13.255Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-11T15:39:59.603Z","emptyReason":null},"description":"当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。 集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。 默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。 支持 wav、mp3、flac、m4a、ogg 等常见 Skill: audio-enhancement-engine Owner: wangminrui2022 Summary: 当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。 集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。 默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。 支持 wav、mp3、flac、m4a、ogg 等常见 Tags: latest:1.0.5 Version history: v1.0.5 | 2026-07-03T","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17de4avhxep6e2yxdr6snee9h83v08q:audio-enhancement-engine","sourceUrl":"https://clawhub.ai/wangminrui2022/audio-enhancement-engine","homepage":"https://clawhub.ai/wangminrui2022/skills/audio-enhancement-engine","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/wangminrui2022/audio-enhancement-engine","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/wangminrui2022/skills/audio-enhancement-engine","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":60,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。 集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T15:39:59.603Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T15:39:59.603Z","emptyReason":null},"stars":null,"forks":null,"downloads":1037,"packageName":null,"latestVersion":"1.0.5","tractionLabel":"1K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T15:39:59.590Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T15:39:59.603Z","lastCrawledAt":"2026-10-11T15:39:59.590Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T15:39:59.590Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.5","createdAt":"2026-07-03T14:11:27.476Z","changelog":"- Major update: Integrated a new audio enhancement backend and expanded functionality. - Added scripts/resemble-enhance-0.0.1 and scripts/resemble_enhance modules for enhanced audio processing workflow. - Updated command-line options to include advanced parameters such as --re, --nfe, --solver, --lambd, and --tau. - Improved flexibility for both VoiceFixer and AudioSR processing modes. - Removed skill-card.md for cleaner project structure.","fileCount":71,"zipByteSize":92951},{"version":"1.0.3","createdAt":"2026-04-20T14:13:54.555Z","changelog":"Version 1.0.3 of audio-enhancement-engine - No file changes detected in this release. - Maintains function and behavior identical to the previous version.","fileCount":12,"zipByteSize":28332},{"version":"1.0.2","createdAt":"2026-04-19T14:44:48.075Z","changelog":"- 技能名称从 audio-enhancement 更新为 audio-enhancement-engine。 - 其余内容无更改。","fileCount":11,"zipByteSize":26554},{"version":"1.0.1","createdAt":"2026-04-19T14:18:56.625Z","changelog":"Version 1.0.1 of audio-enhancement-engine - No file changes detected in this version. - Functionality, user features, and trigger phrases remain unchanged. - No updates to processing procedures or supported formats.","fileCount":11,"zipByteSize":26327},{"version":"1.0.0","createdAt":"2026-04-19T14:10:40.397Z","changelog":"Audio Enhancement Skill 1.0.0 - Initial release providing local audio enhancement and restoration. - Integrates VoiceFixer (general speech enhancement) and AudioSR (high-fidelity super resolution up to 48kHz). - Supports single audio files or entire folders (with batch and recursive processing). - Auto-selects the best enhancement mode (VoiceFixer by default; switches to AudioSR for high-fidelity/music use cases). - Compatible with common formats (wav, mp3, flac, m4a, ogg); all output in high-quality WAV. - Processes only audio files or audio directories—other file types are ignored.","fileCount":11,"zipByteSize":26132}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17de4avhxep6e2yxdr6snee9h83v08q:audio-enhancement-engine","setupComplexity":"low","setupSteps":["Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T17:44:13.253Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-wangminrui2022-audio-enhancement-engine/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-11T15:39:59.603Z","emptyReason":null},"readme":"Skill: audio-enhancement-engine\n\nOwner: wangminrui2022\n\nSummary: 当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\n集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\n默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\n支持 wav、mp3、flac、m4a、ogg 等常见\n\nTags: latest:1.0.5\n\nVersion history:\n\nv1.0.5 | 2026-07-03T14:11:27.476Z | user\n\n- Major update: Integrated a new audio enhancement backend and expanded functionality.\n- Added scripts/resemble-enhance-0.0.1 and scripts/resemble_enhance modules for enhanced audio processing workflow.\n- Updated command-line options to include advanced parameters such as --re, --nfe, --solver, --lambd, and --tau.\n- Improved flexibility for both VoiceFixer and AudioSR processing modes.\n- Removed skill-card.md for cleaner project structure.\n\nv1.0.3 | 2026-04-20T14:13:54.555Z | user\n\nVersion 1.0.3 of audio-enhancement-engine\n\n- No file changes detected in this release.\n- Maintains function and behavior identical to the previous version.\n\nv1.0.2 | 2026-04-19T14:44:48.075Z | user\n\n- 技能名称从 audio-enhancement 更新为 audio-enhancement-engine。\n- 其余内容无更改。\n\nv1.0.1 | 2026-04-19T14:18:56.625Z | user\n\nVersion 1.0.1 of audio-enhancement-engine\n\n- No file changes detected in this version.\n- Functionality, user features, and trigger phrases remain unchanged.\n- No updates to processing procedures or supported formats.\n\nv1.0.0 | 2026-04-19T14:10:40.397Z | user\n\nAudio Enhancement Skill 1.0.0\n\n- Initial release providing local audio enhancement and restoration.\n- Integrates VoiceFixer (general speech enhancement) and AudioSR (high-fidelity super resolution up to 48kHz).\n- Supports single audio files or entire folders (with batch and recursive processing).\n- Auto-selects the best enhancement mode (VoiceFixer by default; switches to AudioSR for high-fidelity/music use cases).\n- Compatible with common formats (wav, mp3, flac, m4a, ogg); all output in high-quality WAV.\n- Processes only audio files or audio directories—other file types are ignored.\n\nArchive index:\n\nArchive v1.0.5: 71 files, 92951 bytes\n\nFiles: README.md (6582b), scripts/config.py (2171b), scripts/enhancer.py (7207b), scripts/ensure_package.py (12283b), scripts/env_manager.py (10520b), scripts/hifi_audio_enhance.py (7695b), scripts/logger_manager.py (2696b), scripts/resemble_enhance/__init__.py (0b), scripts/resemble_enhance/common.py (1802b), scripts/resemble_enhance/data/__init__.py (1267b), scripts/resemble_enhance/data/dataset.py (5592b), scripts/resemble_enhance/data/distorter/__init__.py (33b), scripts/resemble_enhance/data/distorter/base.py (2475b), scripts/resemble_enhance/data/distorter/custom.py (2334b), scripts/resemble_enhance/data/distorter/distorter.py (1161b), scripts/resemble_enhance/data/distorter/sox.py (5132b), scripts/resemble_enhance/data/utils.py (1187b), scripts/resemble_enhance/denoiser/__init__.py (0b), scripts/resemble_enhance/denoiser/__main__.py (1037b), scripts/resemble_enhance/denoiser/denoiser.py (5223b), scripts/resemble_enhance/denoiser/hparams.py (198b), scripts/resemble_enhance/denoiser/inference.py (756b), scripts/resemble_enhance/denoiser/train.py (3558b), scripts/resemble_enhance/denoiser/unet.py (4183b), scripts/resemble_enhance/enhancer/__init__.py (0b), scripts/resemble_enhance/enhancer/__main__.py (3350b), scripts/resemble_enhance/enhancer/download.py (2072b), scripts/resemble_enhance/enhancer/enhancer.py (6822b), scripts/resemble_enhance/enhancer/hparams.py (557b), scripts/resemble_enhance/enhancer/inference.py (1411b), scripts/resemble_enhance/enhancer/lcfm/__init__.py (53b), scripts/resemble_enhance/enhancer/lcfm/cfm.py (10502b), scripts/resemble_enhance/enhancer/lcfm/irmae.py (3872b), scripts/resemble_enhance/enhancer/lcfm/lcfm.py (4336b), scripts/resemble_enhance/enhancer/lcfm/wn.py (3495b), scripts/resemble_enhance/enhancer/train.py (4697b), scripts/resemble_enhance/enhancer/univnet/__init__.py (29b), scripts/resemble_enhance/enhancer/univnet/alias_free_torch/__init__.py (181b), scripts/resemble_enhance/enhancer/univnet/alias_free_torch/filter.py (3457b), scripts/resemble_enhance/enhancer/univnet/alias_free_torch/resample.py (1860b), scripts/resemble_enhance/enhancer/univnet/amp.py (3343b), scripts/resemble_enhance/enhancer/univnet/discriminator.py (6169b), scripts/resemble_enhance/enhancer/univnet/lvcnet.py (10486b), scripts/resemble_enhance/enhancer/univnet/mrstft.py (4324b), scripts/resemble_enhance/enhancer/univnet/univnet.py (2500b), scripts/resemble_enhance/hparams.py (3926b), scripts/resemble_enhance/inference.py (4591b), scripts/resemble_enhance/melspec.py (2056b), scripts/resemble_enhance/utils/__init__.py (215b), scripts/resemble_enhance/utils/control.py (612b), scripts/resemble_enhance/utils/distributed.py (2463b), scripts/resemble_enhance/utils/engine.py (4337b), scripts/resemble_enhance/utils/logging.py (1100b), scripts/resemble_enhance/utils/train_loop.py (8421b), scripts/resemble_enhance/utils/utils.py (1674b), scripts/resemble-enhance-0.0.1/.gitignore (78b), scripts/resemble-enhance-0.0.1/app.py (1569b), scripts/resemble-enhance-0.0.1/config/denoiser.yaml (45b), scripts/resemble-enhance-0.0.1/config/enhancer_stage1.yaml (97b), scripts/resemble-enhance-0.0.1/config/enhancer_stage2.yaml (217b), scripts/resemble-enhance-0.0.1/packages.txt (11b), scripts/resemble-enhance-0.0.1/pyproject.toml (90b), scripts/resemble-enhance-0.0.1/README.md (2487b), scripts/resemble-enhance-0.0.1/requirements.txt (286b), scripts/resemble-enhance-0.0.1/setup.py (1616b), scripts/run_resemble_enhance.py (10444b), scripts/upgrade_torch.py (8814b), scripts/voice_enhance.py (6603b), skill-card.md (2462b), SKILL.md (4983b), _meta.json (143b)\n\nFile v1.0.5:SKILL.md\n\n---\r\nname: audio-enhancement-engine\r\ndescription: |\r\n  当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\r\n  集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\r\n  默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\r\n  支持 wav、mp3、flac、m4a、ogg 等常见格式，完全本地运行，输出统一为高质量 WAV 文件。\r\n\r\n  【重要约束】仅处理音频文件或音频文件夹，其他文件（如视频、图片、文档、纯文本）一律不触发此技能。\r\n\r\n  常见触发口语（越多越好）：\r\n  - “帮我增强这个音频”\r\n  - “修复这个录音的音质”\r\n  - “给这个语音降噪”\r\n  - “把这个音频提升到高保真”\r\n  - “音乐音质增强 这个.mp3”\r\n  - “批量处理音频文件夹”\r\n  - “清理会议录音”\r\n  - “提升音频采样率到48kHz”\r\n  - “语音修复 这个 wav 文件”\r\n  - “高保真增强音频”\r\n  - “老旧录音修复”\r\n  - “音频增强 目录路径”\r\nmetadata:\r\n  openclaw:\r\n    requires:\r\n      bins:\r\n        - python\r\n    user-invocable: true\r\n---\r\n\r\n# Audio Enhancement Skill\r\n\r\n**功能**：本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\r\n\r\n### 触发时机（Triggers）\r\n- 用户提供音频文件（.wav、.mp3、.flac、.m4a、.ogg 等）或音频文件夹路径，并表达增强音质、修复、降噪、高保真等意图。\r\n- 用户说“音频增强”“修复录音”“降噪”“提升音质”“高保真”“48kHz”等关键词。\r\n- 支持单个文件处理或整个文件夹批量处理（支持递归子目录）。\r\n\r\n### 支持的两种增强模式\r\n1. **VoiceFixer 通用语音修复**（默认模式）\r\n   - 擅长语音降噪、提升清晰度、修复轻微失真。\r\n   - 推荐用于：会议录音、访谈、播客、语音笔记、老旧录音。\r\n\r\n2. **AudioSR 高保真音频超级分辨率**（启用 `--hifi` 时）\r\n   - 将音频提升至 48kHz，显著增加高频细节和整体保真度。\r\n   - 推荐用于：音乐、演唱、人声、需要高音质的场景。\r\n\r\n## 参数提取指南\r\n当决定调用此技能时，请从用户消息中准确提取以下参数：\r\n\r\n1. **`<输入路径>`** (必填): 用户提供的音频文件路径或文件夹路径（支持相对/绝对路径）。\r\n2. **`<输出路径>`** (选填): 用户指定的输出文件或目录路径。若未指定，默认在输入同级目录自动添加 `_enhanced` 后缀。\r\n3. **`<模式选择>`** (选填):\r\n   - 默认使用 VoiceFixer。\r\n   - 若用户提到“高保真”“音乐”“48kHz”“超分辨率”等，自动添加 `--hifi` 并使用 AudioSR。\r\n4. **VoiceFixer 专用参数**（默认模式）:\r\n   - `--mode`：0/1/2（推荐 1，默认 1）\r\n   - `--cuda`：是否使用 GPU\r\n   - `-r, --recursive`：是否递归子目录\r\n5. **AudioSR 专用参数**（`--hifi` 模式）:\r\n   - `--model_name`：`basic` 或 `speech`（人声推荐 speech）\r\n   - `--ddim_steps`：扩散步数（默认 50，建议 50-100）\r\n   - `--guidance_scale`：引导尺度（默认 3.5）\r\n   - `--seed`：随机种子（默认 42）\r\n   - `--device`：`cuda` 或 `cpu`\r\n\r\n### 执行步骤\r\n1. **解析路径**：识别用户提供的音频文件或文件夹路径。\r\n2. **模式判断**：根据用户意图判断使用 VoiceFixer（默认）还是 AudioSR（含 `--hifi`）。\r\n3. **默认目标**：若未指定输出路径，默认在输入目录生成带 `_enhanced_48k`（AudioSR）或 `_enhanced`（VoiceFixer）后缀的文件。\r\n4. **调用命令**：使用以下兼容性命令启动脚本（优先 `python3`，失败则 `python`）。脚本会自动检查环境、初始化模型并处理。\r\n\r\n   ```bash\r\n   (python3 scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>] [--re] [--nfe <数值>] [--solver <euler|midpoint|rk4>] [--lambd <数值>] [--tau <数值>]) || (python scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>] [--re] [--nfe <数值>] [--solver <euler|midpoint|rk4>] [--lambd <数值>] [--tau <数值>])\n\nFile v1.0.5:README.md\n\n### audio-enhancement-engine\n\n[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![Python](https://img.shields.io/badge/Python-3.8%2B-green.svg)](https://www.python.org/)\n[![OpenClaw](https://img.shields.io/badge/OpenClaw-Skill-orange.svg)](https://github.com/wangminrui2022)\n\n**audio-enhancement-engine** 本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\n\n---\n\n```markdown\n# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目\n\n```bash\ngit clone https://github.com/wangminrui2022/audio-enhancement-engine.git\n```\n\n> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 2. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）\n\n```bash\n# 修复单个音频文件\npython scripts/enhancer.py -i input/audio.mp3\n\n# 修复整个目录\npython scripts/enhancer.py -i recordings/\n\n# 使用 GPU + 递归处理子目录\npython scripts/enhancer.py -i recordings/ --cuda -r\n```\n\n#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）\n\n```bash\n# 高保真增强单个文件\npython scripts/enhancer.py -i low_quality.wav --hifi\n\n# 高保真增强目录，并指定输出目录\npython scripts/enhancer.py -i music_folder/ -o enhanced_music/ --hifi\n\n# 使用更高参数提升质量\npython scripts/enhancer.py -i input.wav --hifi --ddim_steps 100 --guidance_scale 4.0\n```\n\n#### 方式三：Resemble-Enhance 模型构建的专业级语音增强工具\n\n```bash\n# 高保真增强单个文件\npython scripts/enhancer.py -i low_quality.wav --re\n\n# 高保真增强目录，并指定输出目录\npython scripts/enhancer.py -i music_folder/ -o enhanced_music/ --re\n\n# 使用更高参数提升质量\npython scripts/enhancer.py -i input.wav --re --nfe 64 --solver \"euler\" --lambd 0.75 --tau 0.5\n```\n\n\n---\n### **在 OpenClaw 聊天中**\n\n你可以直接对你的 Agent 说：\n\n帮我增强这个音频 recording.mp3”\n\n复这个会议录音，音质太差了”\n\n给这个音乐提升高保真音质 music.wav”\n\n把这个音频提升到48kHz”\n\n批量处理这个音频文件夹 audio_files/”\n\n高保真增强这个播客音频”\n\n帮我降噪这个语音笔记 voice.m4a”\n\n老旧录音修复 folder_path/”\n\n音乐音质增强 这个.mp3 --hifi”\n\n清理这个失真录音并提升清晰度”\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n### Resemble-Enhance 模型构建的专业级语音增强工具\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--nfe` | int | 64 | 扩散推理步数（越大质量越好，速度越慢） |\n| `--solver` | str | euler | 求解器，可选 euler / midpoint / rk4 |\n| `--lambd` | float | 0.75 | 去噪强度 0~1 |\n| `--tau` | float | 0.5 | 音质保留阈值 |\n\n---\n\n## 📂 项目结构\n\n```\naudio-enhancement-engine/\n├── enhancer.py                # 主程序入口（命令行解析 + 调度）\n├── hifi_audio_enhance.py      # AudioSR 高保真增强模块\n├── voice_enhance.py           # VoiceFixer 语音修复模块\n├── run_resemble_enhance.py    # Resemble-Enhance 语音增强模块\n└── ...\n```\n\n---\n\n## 🎯 使用场景推荐\n\n| 场景 | 推荐模式 | 建议参数 |\n|------|----------|----------|\n| 会议录音、电话录音、播客降噪 | VoiceFixer | `mode=1` + `--cuda` |\n| 老旧磁带、历史录音修复 | VoiceFixer | `mode=1` |\n| 音乐、演唱、人声提升高频细节 | AudioSR | `--hifi --model_name speech` |\n| 需要极致音质（音乐制作） | AudioSR | `--hifi --ddim_steps 100 --guidance_scale 4.0` |\n| 批量处理大量文件 | 任一模式 | 配合 `-o` 指定输出目录 |\n\n---\n\n## ⚠️ 注意事项\n\n1. **首次运行** VoiceFixer 或 AudioSR 时会自动下载模型，请保持网络畅通。\n2. AudioSR 处理速度较慢（尤其是 `ddim_steps` 较大时），建议使用 GPU。\n3. VoiceFixer 默认输出为 `.wav` 格式，AudioSR 默认输出为 48kHz `.wav`。\n4. 建议为不同任务分别创建输出目录，避免混淆。\n\n---\n\n## 🛠️ 未来计划（可选）\n\n- 支持更多音频增强模型\n- 添加图形界面（Gradio / Streamlit）\n- 支持批量配置文件\n- 集成到 OpenClaw 主框架\n- 支持实时音频流处理\n\n---\n\n## 📄 License\n\nApache License\n\n---\n\n## ❤️ 致谢\n\n- [AudioSR](https://github.com/haoheliu/versatile_audio_super_resolution.git) - 音频超级分辨率\n- [VoiceFixer](https://github.com/haoheliu/voicefixer.git) - 语音修复工具\n\n---\n\n**欢迎 Star ⭐ 支持项目！**\n\n如有问题或建议，欢迎提交 Issue 或 Pull Request。\n\n```\n\n---\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/README.md\n\n# Resemble Enhance\n\n[![PyPI](https://img.shields.io/pypi/v/resemble-enhance.svg)](https://pypi.org/project/resemble-enhance/)\n[![Hugging Face Space](https://img.shields.io/badge/Hugging%20Face%20%F0%9F%A4%97-Space-yellow)](https://huggingface.co/spaces/ResembleAI/resemble-enhance)\n[![License](https://img.shields.io/github/license/resemble-ai/Resemble-Enhance.svg)](https://github.com/resemble-ai/resemble-enhance/blob/main/LICENSE)\n\nhttps://github.com/resemble-ai/resemble-enhance/assets/660224/bc3ec943-e795-4646-b119-cce327c810f1\n\nResemble Enhance is an AI-powered tool that aims to improve the overall quality of speech by performing denoising and enhancement. It consists of two modules: a denoiser, which separates speech from a noisy audio, and an enhancer, which further boosts the perceptual audio quality by restoring audio distortions and extending the audio bandwidth. The two models are trained on high-quality 44.1kHz speech data that guarantees the enhancement of your speech with high quality.\n\n## Usage\n\n### Installation\n\n```bash\npip install resemble-enhance\n```\n\n### Enhance\n\n```\nresemble_enhance in_dir out_dir\n```\n\n### Denoise only\n\n```\nresemble_enhance in_dir out_dir --denoise_only\n```\n\n### Web Demo\n\nWe provide a web demo built with Gradio, you can try it out [here](https://huggingface.co/spaces/ResembleAI/resemble-enhance), or also run it locally:\n\n```\npython app.py\n```\n\n## Train your own model\n\n### Data Preparation\n\nYou need to prepare a foreground speech dataset and a background non-speech dataset. In addition, you need to prepare a RIR dataset ([examples](https://github.com/RoyJames/room-impulse-responses)).\n\n```bash\ndata\n├── fg\n│   ├── 00001.wav\n│   └── ...\n├── bg\n│   ├── 00001.wav\n│   └── ...\n└── rir\n    ├── 00001.npy\n    └── ...\n```\n\n### Training\n\n#### Denoiser Warmup\n\nThough the denoiser is trained jointly with the enhancer, it is recommended for a warmup training first.\n\n```bash\npython -m resemble_enhance.denoiser.train --yaml config/denoiser.yaml\n```\n\n#### Enhancer\n\nThen, you can train the enhancer in two stages. The first stage is to train the autoencoder and vocoder. And the second stage is to train the latent conditional flow matching (CFM) model.\n\n##### Stage 1\n\n```bash\npython -m resemble_enhance.enhancer.train --yaml config/enhancer_stage1.yaml\n```\n\n##### Stage 2\n\n```bash\npython -m resemble_enhance.enhancer.train --yaml config/enhancer_stage2.yaml\n```\n\nFile v1.0.5:_meta.json\n\n{\n  \"ownerId\": \"kn77473086ppakmtqf0e4myp8981mhjd\",\n  \"slug\": \"audio-enhancement-engine\",\n  \"version\": \"1.0.5\",\n  \"publishedAt\": 1783087887476\n}\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/config/denoiser.yaml\n\nbatch_size_per_gpu: 32\ntraining_seconds: 3.0\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/config/enhancer_stage1.yaml\n\nlcfm_training_mode: ae\nload_fg_only: true\nbatch_size_per_gpu: 16\ndenoiser_run_dir: runs/denoiser\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/config/enhancer_stage2.yaml\n\nlcfm_training_mode: cfm\nbatch_size_per_gpu: 32\ntraining_seconds: 3.0\ngan_training_start_step: null\nlcfm_z_scale: 6\npraat_augment_prob: 0.2\ndenoiser_run_dir: runs/denoiser\nenhancer_stage1_run_dir: runs/enhancer_stage1\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/packages.txt\n\nlibsox-dev\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/pyproject.toml\n\n[tool.black]\nline-length = 120\ntarget-version = ['py310']\n\n[tool.isort]\nline_length = 120\n\nFile v1.0.5:scripts/resemble-enhance-0.0.1/requirements.txt\n\ncelluloid==0.2.0\ndeepspeed==0.12.4\nlibrosa==0.10.1\nmatplotlib==3.8.1\nnumpy==1.26.2\nomegaconf==2.3.0\npandas==2.1.3\nptflops==0.7.1.2\nrich==13.7.0\nscipy==1.11.4\nsoundfile==0.12.1\ntorch==2.1.1\ntorchaudio==2.1.1\ntorchvision==0.16.1\ntqdm==4.66.1\nresampy==0.4.2\ntabulate==0.8.10\ngradio==4.8.0\n\nFile v1.0.5:skill-card.md\n\n## Description:\n\nAudio Enhancement Engine helps an agent run local audio repair, denoising, and 48 kHz super-resolution workflows for single files or batches using VoiceFixer, AudioSR, or Resemble-Enhance.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[wangminrui2022](https://clawhub.ai/user/wangminrui2022)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nExternal users, developers, and agent operators use this skill to route audio-file enhancement requests into local command-line processing for speech cleanup, music fidelity improvement, and batch WAV output.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can automatically install or upgrade Python packages.\n\nMitigation: Run it in an isolated virtual environment or container that has no access to sensitive files or credentials.\n\nRisk: The skill can download FFmpeg, model files, and large ML dependencies without strong integrity controls.\n\nMitigation: Prefer pinned dependency and model revisions, verify hashes where available, and review downloads before use.\n\nRisk: Audio enhancement workloads may be slow or resource-intensive, especially high-fidelity modes.\n\nMitigation: Test on non-sensitive sample audio first and set resource limits appropriate for the target CPU or GPU environment.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/wangminrui2022/skills/audio-enhancement-engine)\n- [AudioSR project](https://github.com/haoheliu/versatile_audio_super_resolution.git)\n- [VoiceFixer project](https://github.com/haoheliu/voicefixer.git)\n- [Resemble Enhance Hugging Face Space](https://huggingface.co/spaces/ResembleAI/resemble-enhance)\n- [Resemble Enhance package](https://pypi.org/project/resemble-enhance/)\n\n## Skill Output:\n\n**Output Type(s):** [shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline bash command examples and parameter guidance]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Guides local processing of audio files or directories and typically produces high-quality WAV files through the invoked scripts.]\n\n## Skill Version(s):\n\n1.0.5 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.0.3: 12 files, 28332 bytes\n\nFiles: README.md (5628b), scripts/config.py (2171b), scripts/enhancer.py (4015b), scripts/ensure_package.py (12395b), scripts/env_manager.py (10520b), scripts/hifi_audio_enhance.py (7183b), scripts/logger_manager.py (2696b), scripts/upgrade_torch.py (8814b), scripts/voice_enhance.py (6603b), skill-card.md (2107b), SKILL.md (4799b), _meta.json (143b)\n\nFile v1.0.3:SKILL.md\n\n---\r\nname: audio-enhancement-engine\r\ndescription: |\r\n  当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\r\n  集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\r\n  默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\r\n  支持 wav、mp3、flac、m4a、ogg 等常见格式，完全本地运行，输出统一为高质量 WAV 文件。\r\n\r\n  【重要约束】仅处理音频文件或音频文件夹，其他文件（如视频、图片、文档、纯文本）一律不触发此技能。\r\n\r\n  常见触发口语（越多越好）：\r\n  - “帮我增强这个音频”\r\n  - “修复这个录音的音质”\r\n  - “给这个语音降噪”\r\n  - “把这个音频提升到高保真”\r\n  - “音乐音质增强 这个.mp3”\r\n  - “批量处理音频文件夹”\r\n  - “清理会议录音”\r\n  - “提升音频采样率到48kHz”\r\n  - “语音修复 这个 wav 文件”\r\n  - “高保真增强音频”\r\n  - “老旧录音修复”\r\n  - “音频增强 目录路径”\r\nmetadata:\r\n  openclaw:\r\n    requires:\r\n      bins:\r\n        - python\r\n    user-invocable: true\r\n---\r\n\r\n# Audio Enhancement Skill\r\n\r\n**功能**：本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\r\n\r\n### 触发时机（Triggers）\r\n- 用户提供音频文件（.wav、.mp3、.flac、.m4a、.ogg 等）或音频文件夹路径，并表达增强音质、修复、降噪、高保真等意图。\r\n- 用户说“音频增强”“修复录音”“降噪”“提升音质”“高保真”“48kHz”等关键词。\r\n- 支持单个文件处理或整个文件夹批量处理（支持递归子目录）。\r\n\r\n### 支持的两种增强模式\r\n1. **VoiceFixer 通用语音修复**（默认模式）\r\n   - 擅长语音降噪、提升清晰度、修复轻微失真。\r\n   - 推荐用于：会议录音、访谈、播客、语音笔记、老旧录音。\r\n\r\n2. **AudioSR 高保真音频超级分辨率**（启用 `--hifi` 时）\r\n   - 将音频提升至 48kHz，显著增加高频细节和整体保真度。\r\n   - 推荐用于：音乐、演唱、人声、需要高音质的场景。\r\n\r\n## 参数提取指南\r\n当决定调用此技能时，请从用户消息中准确提取以下参数：\r\n\r\n1. **`<输入路径>`** (必填): 用户提供的音频文件路径或文件夹路径（支持相对/绝对路径）。\r\n2. **`<输出路径>`** (选填): 用户指定的输出文件或目录路径。若未指定，默认在输入同级目录自动添加 `_enhanced` 后缀。\r\n3. **`<模式选择>`** (选填):\r\n   - 默认使用 VoiceFixer。\r\n   - 若用户提到“高保真”“音乐”“48kHz”“超分辨率”等，自动添加 `--hifi` 并使用 AudioSR。\r\n4. **VoiceFixer 专用参数**（默认模式）:\r\n   - `--mode`：0/1/2（推荐 1，默认 1）\r\n   - `--cuda`：是否使用 GPU\r\n   - `-r, --recursive`：是否递归子目录\r\n5. **AudioSR 专用参数**（`--hifi` 模式）:\r\n   - `--model_name`：`basic` 或 `speech`（人声推荐 speech）\r\n   - `--ddim_steps`：扩散步数（默认 50，建议 50-100）\r\n   - `--guidance_scale`：引导尺度（默认 3.5）\r\n   - `--seed`：随机种子（默认 42）\r\n   - `--device`：`cuda` 或 `cpu`\r\n\r\n### 执行步骤\r\n1. **解析路径**：识别用户提供的音频文件或文件夹路径。\r\n2. **模式判断**：根据用户意图判断使用 VoiceFixer（默认）还是 AudioSR（含 `--hifi`）。\r\n3. **默认目标**：若未指定输出路径，默认在输入目录生成带 `_enhanced_48k`（AudioSR）或 `_enhanced`（VoiceFixer）后缀的文件。\r\n4. **调用命令**：使用以下兼容性命令启动脚本（优先 `python3`，失败则 `python`）。脚本会自动检查环境、初始化模型并处理。\r\n\r\n   ```bash\r\n   (python3 scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>]) || (python scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>])\n\nFile v1.0.3:README.md\n\n### audio-enhancement-engine\n\n[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![Python](https://img.shields.io/badge/Python-3.8%2B-green.svg)](https://www.python.org/)\n[![OpenClaw](https://img.shields.io/badge/OpenClaw-Skill-orange.svg)](https://github.com/wangminrui2022)\n\n**audio-enhancement-engine** 本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\n\n---\n\n```markdown\n# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目\n\n```bash\ngit clone https://github.com/wangminrui2022/audio-enhancement-engine.git\n```\n\n> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 2. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）\n\n```bash\n# 修复单个音频文件\npython main.py -i input/audio.mp3\n\n# 修复整个目录\npython main.py -i recordings/\n\n# 使用 GPU + 递归处理子目录\npython main.py -i recordings/ --cuda -r\n```\n\n#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）\n\n```bash\n# 高保真增强单个文件\npython main.py -i low_quality.wav --hifi\n\n# 高保真增强目录，并指定输出目录\npython main.py -i music_folder/ -o enhanced_music/ --hifi\n\n# 使用更高参数提升质量\npython main.py -i input.wav --hifi --ddim_steps 100 --guidance_scale 4.0\n```\n\n---\n### **在 OpenClaw 聊天中**\n\n你可以直接对你的 Agent 说：\n\n帮我增强这个音频 recording.mp3”\n\n复这个会议录音，音质太差了”\n\n给这个音乐提升高保真音质 music.wav”\n\n把这个音频提升到48kHz”\n\n批量处理这个音频文件夹 audio_files/”\n\n高保真增强这个播客音频”\n\n帮我降噪这个语音笔记 voice.m4a”\n\n老旧录音修复 folder_path/”\n\n音乐音质增强 这个.mp3 --hifi”\n\n清理这个失真录音并提升清晰度”\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n---\n\n## 📂 项目结构\n\n```\naudio-enhancement-engine/\n├── enhancer.py                # 主程序入口（命令行解析 + 调度）\n├── hifi_audio_enhance.py      # AudioSR 高保真增强模块\n├── voice_enhance.py           # VoiceFixer 语音修复模块\n└── ...\n```\n\n---\n\n## 🎯 使用场景推荐\n\n| 场景 | 推荐模式 | 建议参数 |\n|------|----------|----------|\n| 会议录音、电话录音、播客降噪 | VoiceFixer | `mode=1` + `--cuda` |\n| 老旧磁带、历史录音修复 | VoiceFixer | `mode=1` |\n| 音乐、演唱、人声提升高频细节 | AudioSR | `--hifi --model_name speech` |\n| 需要极致音质（音乐制作） | AudioSR | `--hifi --ddim_steps 100 --guidance_scale 4.0` |\n| 批量处理大量文件 | 任一模式 | 配合 `-o` 指定输出目录 |\n\n---\n\n## ⚠️ 注意事项\n\n1. **首次运行** VoiceFixer 或 AudioSR 时会自动下载模型，请保持网络畅通。\n2. AudioSR 处理速度较慢（尤其是 `ddim_steps` 较大时），建议使用 GPU。\n3. VoiceFixer 默认输出为 `.wav` 格式，AudioSR 默认输出为 48kHz `.wav`。\n4. 建议为不同任务分别创建输出目录，避免混淆。\n\n---\n\n## 🛠️ 未来计划（可选）\n\n- 支持更多音频增强模型\n- 添加图形界面（Gradio / Streamlit）\n- 支持批量配置文件\n- 集成到 OpenClaw 主框架\n- 支持实时音频流处理\n\n---\n\n## 📄 License\n\nApache License\n\n---\n\n## ❤️ 致谢\n\n- [AudioSR](https://github.com/haoheliu/versatile_audio_super_resolution.git) - 音频超级分辨率\n- [VoiceFixer](https://github.com/haoheliu/voicefixer.git) - 语音修复工具\n\n---\n\n**欢迎 Star ⭐ 支持项目！**\n\n如有问题或建议，欢迎提交 Issue 或 Pull Request。\n\n```\n\n---\n\nFile v1.0.3:_meta.json\n\n{\n  \"ownerId\": \"kn77473086ppakmtqf0e4myp8981mhjd\",\n  \"slug\": \"audio-enhancement-engine\",\n  \"version\": \"1.0.3\",\n  \"publishedAt\": 1776694434555\n}\n\nFile v1.0.3:skill-card.md\n\n## Description: <br>\nAudio Enhancement Engine helps agents run local audio enhancement and repair for audio files or folders using VoiceFixer for speech restoration and AudioSR for 48 kHz high-fidelity super-resolution. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[wangminrui2022](https://clawhub.ai/user/wangminrui2022) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nExternal users, developers, and audio workflows use this skill to enhance speech recordings, meetings, podcasts, music, and batches of common audio formats locally. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: First run can download dependencies or model code and change the Python environment. <br>\nMitigation: Run in an isolated container or disposable virtual environment, and review and pin dependency sources before use. <br>\nRisk: Recursive batch mode can process more local audio files than intended. <br>\nMitigation: Point the skill only at audio files or folders explicitly intended for enhancement and review output paths before running. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/wangminrui2022/audio-enhancement-engine) <br>\n- [AudioSR](https://github.com/haoheliu/versatile_audio_super_resolution.git) <br>\n- [VoiceFixer](https://github.com/haoheliu/voicefixer.git) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Shell commands, Files, Guidance] <br>\n**Output Format:** [Markdown with inline bash commands and generated WAV audio files] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Enhanced audio is written as WAV; VoiceFixer uses an _enhanced suffix and AudioSR produces 48 kHz output.] <br>\n\n## Skill Version(s): <br>\n1.0.3 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.2: 11 files, 26554 bytes\n\nFiles: README.md (5628b), scripts/config.py (2171b), scripts/enhancer.py (3624b), scripts/ensure_package.py (10769b), scripts/env_manager.py (10520b), scripts/hifi_audio_enhance.py (7183b), scripts/logger_manager.py (2696b), scripts/upgrade_torch.py (8814b), scripts/voice_enhance.py (6718b), SKILL.md (4799b), _meta.json (143b)\n\nFile v1.0.2:SKILL.md\n\n---\r\nname: audio-enhancement-engine\r\ndescription: |\r\n  当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\r\n  集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\r\n  默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\r\n  支持 wav、mp3、flac、m4a、ogg 等常见格式，完全本地运行，输出统一为高质量 WAV 文件。\r\n\r\n  【重要约束】仅处理音频文件或音频文件夹，其他文件（如视频、图片、文档、纯文本）一律不触发此技能。\r\n\r\n  常见触发口语（越多越好）：\r\n  - “帮我增强这个音频”\r\n  - “修复这个录音的音质”\r\n  - “给这个语音降噪”\r\n  - “把这个音频提升到高保真”\r\n  - “音乐音质增强 这个.mp3”\r\n  - “批量处理音频文件夹”\r\n  - “清理会议录音”\r\n  - “提升音频采样率到48kHz”\r\n  - “语音修复 这个 wav 文件”\r\n  - “高保真增强音频”\r\n  - “老旧录音修复”\r\n  - “音频增强 目录路径”\r\nmetadata:\r\n  openclaw:\r\n    requires:\r\n      bins:\r\n        - python\r\n    user-invocable: true\r\n---\r\n\r\n# Audio Enhancement Skill\r\n\r\n**功能**：本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\r\n\r\n### 触发时机（Triggers）\r\n- 用户提供音频文件（.wav、.mp3、.flac、.m4a、.ogg 等）或音频文件夹路径，并表达增强音质、修复、降噪、高保真等意图。\r\n- 用户说“音频增强”“修复录音”“降噪”“提升音质”“高保真”“48kHz”等关键词。\r\n- 支持单个文件处理或整个文件夹批量处理（支持递归子目录）。\r\n\r\n### 支持的两种增强模式\r\n1. **VoiceFixer 通用语音修复**（默认模式）\r\n   - 擅长语音降噪、提升清晰度、修复轻微失真。\r\n   - 推荐用于：会议录音、访谈、播客、语音笔记、老旧录音。\r\n\r\n2. **AudioSR 高保真音频超级分辨率**（启用 `--hifi` 时）\r\n   - 将音频提升至 48kHz，显著增加高频细节和整体保真度。\r\n   - 推荐用于：音乐、演唱、人声、需要高音质的场景。\r\n\r\n## 参数提取指南\r\n当决定调用此技能时，请从用户消息中准确提取以下参数：\r\n\r\n1. **`<输入路径>`** (必填): 用户提供的音频文件路径或文件夹路径（支持相对/绝对路径）。\r\n2. **`<输出路径>`** (选填): 用户指定的输出文件或目录路径。若未指定，默认在输入同级目录自动添加 `_enhanced` 后缀。\r\n3. **`<模式选择>`** (选填):\r\n   - 默认使用 VoiceFixer。\r\n   - 若用户提到“高保真”“音乐”“48kHz”“超分辨率”等，自动添加 `--hifi` 并使用 AudioSR。\r\n4. **VoiceFixer 专用参数**（默认模式）:\r\n   - `--mode`：0/1/2（推荐 1，默认 1）\r\n   - `--cuda`：是否使用 GPU\r\n   - `-r, --recursive`：是否递归子目录\r\n5. **AudioSR 专用参数**（`--hifi` 模式）:\r\n   - `--model_name`：`basic` 或 `speech`（人声推荐 speech）\r\n   - `--ddim_steps`：扩散步数（默认 50，建议 50-100）\r\n   - `--guidance_scale`：引导尺度（默认 3.5）\r\n   - `--seed`：随机种子（默认 42）\r\n   - `--device`：`cuda` 或 `cpu`\r\n\r\n### 执行步骤\r\n1. **解析路径**：识别用户提供的音频文件或文件夹路径。\r\n2. **模式判断**：根据用户意图判断使用 VoiceFixer（默认）还是 AudioSR（含 `--hifi`）。\r\n3. **默认目标**：若未指定输出路径，默认在输入目录生成带 `_enhanced_48k`（AudioSR）或 `_enhanced`（VoiceFixer）后缀的文件。\r\n4. **调用命令**：使用以下兼容性命令启动脚本（优先 `python3`，失败则 `python`）。脚本会自动检查环境、初始化模型并处理。\r\n\r\n   ```bash\r\n   (python3 scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>]) || (python scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>])\n\nFile v1.0.2:README.md\n\n### audio-enhancement-engine\n\n[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![Python](https://img.shields.io/badge/Python-3.8%2B-green.svg)](https://www.python.org/)\n[![OpenClaw](https://img.shields.io/badge/OpenClaw-Skill-orange.svg)](https://github.com/wangminrui2022)\n\n**audio-enhancement-engine** 本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\n\n---\n\n```markdown\n# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目\n\n```bash\ngit clone https://github.com/wangminrui2022/audio-enhancement-engine.git\n```\n\n> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 2. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）\n\n```bash\n# 修复单个音频文件\npython main.py -i input/audio.mp3\n\n# 修复整个目录\npython main.py -i recordings/\n\n# 使用 GPU + 递归处理子目录\npython main.py -i recordings/ --cuda -r\n```\n\n#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）\n\n```bash\n# 高保真增强单个文件\npython main.py -i low_quality.wav --hifi\n\n# 高保真增强目录，并指定输出目录\npython main.py -i music_folder/ -o enhanced_music/ --hifi\n\n# 使用更高参数提升质量\npython main.py -i input.wav --hifi --ddim_steps 100 --guidance_scale 4.0\n```\n\n---\n### **在 OpenClaw 聊天中**\n\n你可以直接对你的 Agent 说：\n\n帮我增强这个音频 recording.mp3”\n\n复这个会议录音，音质太差了”\n\n给这个音乐提升高保真音质 music.wav”\n\n把这个音频提升到48kHz”\n\n批量处理这个音频文件夹 audio_files/”\n\n高保真增强这个播客音频”\n\n帮我降噪这个语音笔记 voice.m4a”\n\n老旧录音修复 folder_path/”\n\n音乐音质增强 这个.mp3 --hifi”\n\n清理这个失真录音并提升清晰度”\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n---\n\n## 📂 项目结构\n\n```\naudio-enhancement-engine/\n├── enhancer.py                # 主程序入口（命令行解析 + 调度）\n├── hifi_audio_enhance.py      # AudioSR 高保真增强模块\n├── voice_enhance.py           # VoiceFixer 语音修复模块\n└── ...\n```\n\n---\n\n## 🎯 使用场景推荐\n\n| 场景 | 推荐模式 | 建议参数 |\n|------|----------|----------|\n| 会议录音、电话录音、播客降噪 | VoiceFixer | `mode=1` + `--cuda` |\n| 老旧磁带、历史录音修复 | VoiceFixer | `mode=1` |\n| 音乐、演唱、人声提升高频细节 | AudioSR | `--hifi --model_name speech` |\n| 需要极致音质（音乐制作） | AudioSR | `--hifi --ddim_steps 100 --guidance_scale 4.0` |\n| 批量处理大量文件 | 任一模式 | 配合 `-o` 指定输出目录 |\n\n---\n\n## ⚠️ 注意事项\n\n1. **首次运行** VoiceFixer 或 AudioSR 时会自动下载模型，请保持网络畅通。\n2. AudioSR 处理速度较慢（尤其是 `ddim_steps` 较大时），建议使用 GPU。\n3. VoiceFixer 默认输出为 `.wav` 格式，AudioSR 默认输出为 48kHz `.wav`。\n4. 建议为不同任务分别创建输出目录，避免混淆。\n\n---\n\n## 🛠️ 未来计划（可选）\n\n- 支持更多音频增强模型\n- 添加图形界面（Gradio / Streamlit）\n- 支持批量配置文件\n- 集成到 OpenClaw 主框架\n- 支持实时音频流处理\n\n---\n\n## 📄 License\n\nApache License\n\n---\n\n## ❤️ 致谢\n\n- [AudioSR](https://github.com/haoheliu/versatile_audio_super_resolution.git) - 音频超级分辨率\n- [VoiceFixer](https://github.com/haoheliu/voicefixer.git) - 语音修复工具\n\n---\n\n**欢迎 Star ⭐ 支持项目！**\n\n如有问题或建议，欢迎提交 Issue 或 Pull Request。\n\n```\n\n---\n\nFile v1.0.2:_meta.json\n\n{\n  \"ownerId\": \"kn77473086ppakmtqf0e4myp8981mhjd\",\n  \"slug\": \"audio-enhancement-engine\",\n  \"version\": \"1.0.2\",\n  \"publishedAt\": 1776609888075\n}\n\nArchive v1.0.1: 11 files, 26327 bytes\n\nFiles: README.md (5008b), scripts/config.py (2171b), scripts/enhancer.py (3624b), scripts/ensure_package.py (10769b), scripts/env_manager.py (10520b), scripts/hifi_audio_enhance.py (7183b), scripts/logger_manager.py (2696b), scripts/upgrade_torch.py (8814b), scripts/voice_enhance.py (6718b), SKILL.md (4792b), _meta.json (143b)\n\nFile v1.0.1:SKILL.md\n\n---\r\nname: audio-enhancement\r\ndescription: |\r\n  当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\r\n  集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\r\n  默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\r\n  支持 wav、mp3、flac、m4a、ogg 等常见格式，完全本地运行，输出统一为高质量 WAV 文件。\r\n\r\n  【重要约束】仅处理音频文件或音频文件夹，其他文件（如视频、图片、文档、纯文本）一律不触发此技能。\r\n\r\n  常见触发口语（越多越好）：\r\n  - “帮我增强这个音频”\r\n  - “修复这个录音的音质”\r\n  - “给这个语音降噪”\r\n  - “把这个音频提升到高保真”\r\n  - “音乐音质增强 这个.mp3”\r\n  - “批量处理音频文件夹”\r\n  - “清理会议录音”\r\n  - “提升音频采样率到48kHz”\r\n  - “语音修复 这个 wav 文件”\r\n  - “高保真增强音频”\r\n  - “老旧录音修复”\r\n  - “音频增强 目录路径”\r\nmetadata:\r\n  openclaw:\r\n    requires:\r\n      bins:\r\n        - python\r\n    user-invocable: true\r\n---\r\n\r\n# Audio Enhancement Skill\r\n\r\n**功能**：本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\r\n\r\n### 触发时机（Triggers）\r\n- 用户提供音频文件（.wav、.mp3、.flac、.m4a、.ogg 等）或音频文件夹路径，并表达增强音质、修复、降噪、高保真等意图。\r\n- 用户说“音频增强”“修复录音”“降噪”“提升音质”“高保真”“48kHz”等关键词。\r\n- 支持单个文件处理或整个文件夹批量处理（支持递归子目录）。\r\n\r\n### 支持的两种增强模式\r\n1. **VoiceFixer 通用语音修复**（默认模式）\r\n   - 擅长语音降噪、提升清晰度、修复轻微失真。\r\n   - 推荐用于：会议录音、访谈、播客、语音笔记、老旧录音。\r\n\r\n2. **AudioSR 高保真音频超级分辨率**（启用 `--hifi` 时）\r\n   - 将音频提升至 48kHz，显著增加高频细节和整体保真度。\r\n   - 推荐用于：音乐、演唱、人声、需要高音质的场景。\r\n\r\n## 参数提取指南\r\n当决定调用此技能时，请从用户消息中准确提取以下参数：\r\n\r\n1. **`<输入路径>`** (必填): 用户提供的音频文件路径或文件夹路径（支持相对/绝对路径）。\r\n2. **`<输出路径>`** (选填): 用户指定的输出文件或目录路径。若未指定，默认在输入同级目录自动添加 `_enhanced` 后缀。\r\n3. **`<模式选择>`** (选填):\r\n   - 默认使用 VoiceFixer。\r\n   - 若用户提到“高保真”“音乐”“48kHz”“超分辨率”等，自动添加 `--hifi` 并使用 AudioSR。\r\n4. **VoiceFixer 专用参数**（默认模式）:\r\n   - `--mode`：0/1/2（推荐 1，默认 1）\r\n   - `--cuda`：是否使用 GPU\r\n   - `-r, --recursive`：是否递归子目录\r\n5. **AudioSR 专用参数**（`--hifi` 模式）:\r\n   - `--model_name`：`basic` 或 `speech`（人声推荐 speech）\r\n   - `--ddim_steps`：扩散步数（默认 50，建议 50-100）\r\n   - `--guidance_scale`：引导尺度（默认 3.5）\r\n   - `--seed`：随机种子（默认 42）\r\n   - `--device`：`cuda` 或 `cpu`\r\n\r\n### 执行步骤\r\n1. **解析路径**：识别用户提供的音频文件或文件夹路径。\r\n2. **模式判断**：根据用户意图判断使用 VoiceFixer（默认）还是 AudioSR（含 `--hifi`）。\r\n3. **默认目标**：若未指定输出路径，默认在输入目录生成带 `_enhanced_48k`（AudioSR）或 `_enhanced`（VoiceFixer）后缀的文件。\r\n4. **调用命令**：使用以下兼容性命令启动脚本（优先 `python3`，失败则 `python`）。脚本会自动检查环境、初始化模型并处理。\r\n\r\n   ```bash\r\n   (python3 scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>]) || (python scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>])\n\nFile v1.0.1:README.md\n\n---\n\n```markdown\n# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目\n\n```bash\ngit clone https://github.com/wangminrui2022/audio-enhancement-engine.git\n```\n\n> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 2. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）\n\n```bash\n# 修复单个音频文件\npython main.py -i input/audio.mp3\n\n# 修复整个目录\npython main.py -i recordings/\n\n# 使用 GPU + 递归处理子目录\npython main.py -i recordings/ --cuda -r\n```\n\n#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）\n\n```bash\n# 高保真增强单个文件\npython main.py -i low_quality.wav --hifi\n\n# 高保真增强目录，并指定输出目录\npython main.py -i music_folder/ -o enhanced_music/ --hifi\n\n# 使用更高参数提升质量\npython main.py -i input.wav --hifi --ddim_steps 100 --guidance_scale 4.0\n```\n\n---\n### **在 OpenClaw 聊天中**\n\n你可以直接对你的 Agent 说：\n\n帮我增强这个音频 recording.mp3”\n\n复这个会议录音，音质太差了”\n\n给这个音乐提升高保真音质 music.wav”\n\n把这个音频提升到48kHz”\n\n批量处理这个音频文件夹 audio_files/”\n\n高保真增强这个播客音频”\n\n帮我降噪这个语音笔记 voice.m4a”\n\n老旧录音修复 folder_path/”\n\n音乐音质增强 这个.mp3 --hifi”\n\n清理这个失真录音并提升清晰度”\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n---\n\n## 📂 项目结构\n\n```\nOpenClaw-Audio-Skill/\n├── main.py                    # 主程序入口（命令行解析 + 调度）\n├── hifi_audio_enhance.py      # AudioSR 高保真增强模块\n├── voice_enhance.py           # VoiceFixer 语音修复模块\n├── requirements.txt\n├── README.md\n└── ...\n```\n\n---\n\n## 🎯 使用场景推荐\n\n| 场景 | 推荐模式 | 建议参数 |\n|------|----------|----------|\n| 会议录音、电话录音、播客降噪 | VoiceFixer | `mode=1` + `--cuda` |\n| 老旧磁带、历史录音修复 | VoiceFixer | `mode=1` |\n| 音乐、演唱、人声提升高频细节 | AudioSR | `--hifi --model_name speech` |\n| 需要极致音质（音乐制作） | AudioSR | `--hifi --ddim_steps 100 --guidance_scale 4.0` |\n| 批量处理大量文件 | 任一模式 | 配合 `-o` 指定输出目录 |\n\n---\n\n## ⚠️ 注意事项\n\n1. **首次运行** VoiceFixer 或 AudioSR 时会自动下载模型，请保持网络畅通。\n2. AudioSR 处理速度较慢（尤其是 `ddim_steps` 较大时），建议使用 GPU。\n3. VoiceFixer 默认输出为 `.wav` 格式，AudioSR 默认输出为 48kHz `.wav`。\n4. 建议为不同任务分别创建输出目录，避免混淆。\n\n---\n\n## 🛠️ 未来计划（可选）\n\n- 支持更多音频增强模型\n- 添加图形界面（Gradio / Streamlit）\n- 支持批量配置文件\n- 集成到 OpenClaw 主框架\n- 支持实时音频流处理\n\n---\n\n## 📄 License\n\nMIT License\n\n---\n\n## ❤️ 致谢\n\n- [AudioSR](https://github.com/haoheliu/audiosr) - 音频超级分辨率\n- [VoiceFixer](https://github.com/haoheliu/voicefixer) - 语音修复工具\n\n---\n\n**欢迎 Star ⭐ 支持项目！**\n\n如有问题或建议，欢迎提交 Issue 或 Pull Request。\n\n```\n\n---\n\nFile v1.0.1:_meta.json\n\n{\n  \"ownerId\": \"kn77473086ppakmtqf0e4myp8981mhjd\",\n  \"slug\": \"audio-enhancement-engine\",\n  \"version\": \"1.0.1\",\n  \"publishedAt\": 1776608336625\n}\n\nArchive v1.0.0: 11 files, 26132 bytes\n\nFiles: README.md (4595b), scripts/config.py (2171b), scripts/enhancer.py (3624b), scripts/ensure_package.py (10769b), scripts/env_manager.py (10520b), scripts/hifi_audio_enhance.py (7183b), scripts/logger_manager.py (2696b), scripts/upgrade_torch.py (8814b), scripts/voice_enhance.py (6718b), SKILL.md (4792b), _meta.json (143b)\n\nFile v1.0.0:SKILL.md\n\n---\r\nname: audio-enhancement\r\ndescription: |\r\n  当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\r\n  集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\r\n  默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\r\n  支持 wav、mp3、flac、m4a、ogg 等常见格式，完全本地运行，输出统一为高质量 WAV 文件。\r\n\r\n  【重要约束】仅处理音频文件或音频文件夹，其他文件（如视频、图片、文档、纯文本）一律不触发此技能。\r\n\r\n  常见触发口语（越多越好）：\r\n  - “帮我增强这个音频”\r\n  - “修复这个录音的音质”\r\n  - “给这个语音降噪”\r\n  - “把这个音频提升到高保真”\r\n  - “音乐音质增强 这个.mp3”\r\n  - “批量处理音频文件夹”\r\n  - “清理会议录音”\r\n  - “提升音频采样率到48kHz”\r\n  - “语音修复 这个 wav 文件”\r\n  - “高保真增强音频”\r\n  - “老旧录音修复”\r\n  - “音频增强 目录路径”\r\nmetadata:\r\n  openclaw:\r\n    requires:\r\n      bins:\r\n        - python\r\n    user-invocable: true\r\n---\r\n\r\n# Audio Enhancement Skill\r\n\r\n**功能**：本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\r\n\r\n### 触发时机（Triggers）\r\n- 用户提供音频文件（.wav、.mp3、.flac、.m4a、.ogg 等）或音频文件夹路径，并表达增强音质、修复、降噪、高保真等意图。\r\n- 用户说“音频增强”“修复录音”“降噪”“提升音质”“高保真”“48kHz”等关键词。\r\n- 支持单个文件处理或整个文件夹批量处理（支持递归子目录）。\r\n\r\n### 支持的两种增强模式\r\n1. **VoiceFixer 通用语音修复**（默认模式）\r\n   - 擅长语音降噪、提升清晰度、修复轻微失真。\r\n   - 推荐用于：会议录音、访谈、播客、语音笔记、老旧录音。\r\n\r\n2. **AudioSR 高保真音频超级分辨率**（启用 `--hifi` 时）\r\n   - 将音频提升至 48kHz，显著增加高频细节和整体保真度。\r\n   - 推荐用于：音乐、演唱、人声、需要高音质的场景。\r\n\r\n## 参数提取指南\r\n当决定调用此技能时，请从用户消息中准确提取以下参数：\r\n\r\n1. **`<输入路径>`** (必填): 用户提供的音频文件路径或文件夹路径（支持相对/绝对路径）。\r\n2. **`<输出路径>`** (选填): 用户指定的输出文件或目录路径。若未指定，默认在输入同级目录自动添加 `_enhanced` 后缀。\r\n3. **`<模式选择>`** (选填):\r\n   - 默认使用 VoiceFixer。\r\n   - 若用户提到“高保真”“音乐”“48kHz”“超分辨率”等，自动添加 `--hifi` 并使用 AudioSR。\r\n4. **VoiceFixer 专用参数**（默认模式）:\r\n   - `--mode`：0/1/2（推荐 1，默认 1）\r\n   - `--cuda`：是否使用 GPU\r\n   - `-r, --recursive`：是否递归子目录\r\n5. **AudioSR 专用参数**（`--hifi` 模式）:\r\n   - `--model_name`：`basic` 或 `speech`（人声推荐 speech）\r\n   - `--ddim_steps`：扩散步数（默认 50，建议 50-100）\r\n   - `--guidance_scale`：引导尺度（默认 3.5）\r\n   - `--seed`：随机种子（默认 42）\r\n   - `--device`：`cuda` 或 `cpu`\r\n\r\n### 执行步骤\r\n1. **解析路径**：识别用户提供的音频文件或文件夹路径。\r\n2. **模式判断**：根据用户意图判断使用 VoiceFixer（默认）还是 AudioSR（含 `--hifi`）。\r\n3. **默认目标**：若未指定输出路径，默认在输入目录生成带 `_enhanced_48k`（AudioSR）或 `_enhanced`（VoiceFixer）后缀的文件。\r\n4. **调用命令**：使用以下兼容性命令启动脚本（优先 `python3`，失败则 `python`）。脚本会自动检查环境、初始化模型并处理。\r\n\r\n   ```bash\r\n   (python3 scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>]) || (python scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>])\n\nFile v1.0.0:README.md\n\n---\n\n```markdown\n# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目\n\n```bash\ngit clone https://github.com/你的用户名/OpenClaw-Audio-Skill.git\ncd OpenClaw-Audio-Skill\n```\n\n### 2. 安装依赖\n\n```bash\npip install -r requirements.txt\n```\n\n> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 3. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）\n\n```bash\n# 修复单个音频文件\npython main.py -i input/audio.mp3\n\n# 修复整个目录\npython main.py -i recordings/\n\n# 使用 GPU + 递归处理子目录\npython main.py -i recordings/ --cuda -r\n```\n\n#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）\n\n```bash\n# 高保真增强单个文件\npython main.py -i low_quality.wav --hifi\n\n# 高保真增强目录，并指定输出目录\npython main.py -i music_folder/ -o enhanced_music/ --hifi\n\n# 使用更高参数提升质量\npython main.py -i input.wav --hifi --ddim_steps 100 --guidance_scale 4.0\n```\n\n---\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n---\n\n## 📂 项目结构\n\n```\nOpenClaw-Audio-Skill/\n├── main.py                    # 主程序入口（命令行解析 + 调度）\n├── hifi_audio_enhance.py      # AudioSR 高保真增强模块\n├── voice_enhance.py           # VoiceFixer 语音修复模块\n├── requirements.txt\n├── README.md\n└── ...\n```\n\n---\n\n## 🎯 使用场景推荐\n\n| 场景 | 推荐模式 | 建议参数 |\n|------|----------|----------|\n| 会议录音、电话录音、播客降噪 | VoiceFixer | `mode=1` + `--cuda` |\n| 老旧磁带、历史录音修复 | VoiceFixer | `mode=1` |\n| 音乐、演唱、人声提升高频细节 | AudioSR | `--hifi --model_name speech` |\n| 需要极致音质（音乐制作） | AudioSR | `--hifi --ddim_steps 100 --guidance_scale 4.0` |\n| 批量处理大量文件 | 任一模式 | 配合 `-o` 指定输出目录 |\n\n---\n\n## ⚠️ 注意事项\n\n1. **首次运行** VoiceFixer 或 AudioSR 时会自动下载模型，请保持网络畅通。\n2. AudioSR 处理速度较慢（尤其是 `ddim_steps` 较大时），建议使用 GPU。\n3. VoiceFixer 默认输出为 `.wav` 格式，AudioSR 默认输出为 48kHz `.wav`。\n4. 建议为不同任务分别创建输出目录，避免混淆。\n\n---\n\n## 🛠️ 未来计划（可选）\n\n- 支持更多音频增强模型\n- 添加图形界面（Gradio / Streamlit）\n- 支持批量配置文件\n- 集成到 OpenClaw 主框架\n- 支持实时音频流处理\n\n---\n\n## 📄 License\n\nMIT License\n\n---\n\n## ❤️ 致谢\n\n- [AudioSR](https://github.com/haoheliu/audiosr) - 音频超级分辨率\n- [VoiceFixer](https://github.com/haoheliu/voicefixer) - 语音修复工具\n\n---\n\n**欢迎 Star ⭐ 支持项目！**\n\n如有问题或建议，欢迎提交 Issue 或 Pull Request。\n\n```\n\n---\n\nFile v1.0.0:_meta.json\n\n{\n  \"ownerId\": \"kn77473086ppakmtqf0e4myp8981mhjd\",\n  \"slug\": \"audio-enhancement-engine\",\n  \"version\": \"1.0.0\",\n  \"publishedAt\": 1776607840397\n}","readmeExcerpt":"Skill: audio-enhancement-engine Owner: wangminrui2022 Summary: 当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。 集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。 默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。 支持 wav、mp3、flac、m4a、ogg 等常见 Tags: latest:1.0.5 Version history: v1.0.5 | 2026-07-03T","codeSnippets":[],"executableExamples":[{"language":"markdown","snippet":"# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目"},{"language":"text","snippet":"> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 2. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）"},{"language":"text","snippet":"#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）"},{"language":"text","snippet":"#### 方式三：Resemble-Enhance 模型构建的专业级语音增强工具"},{"language":"text","snippet":"---\n### **在 OpenClaw 聊天中**\n\n你可以直接对你的 Agent 说：\n\n帮我增强这个音频 recording.mp3”\n\n复这个会议录音，音质太差了”\n\n给这个音乐提升高保真音质 music.wav”\n\n把这个音频提升到48kHz”\n\n批量处理这个音频文件夹 audio_files/”\n\n高保真增强这个播客音频”\n\n帮我降噪这个语音笔记 voice.m4a”\n\n老旧录音修复 folder_path/”\n\n音乐音质增强 这个.mp3 --hifi”\n\n清理这个失真录音并提升清晰度”\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n### Resemble-Enhance 模型构建的专业级语音增强工具\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--nfe` | int | 64 | 扩散推理步数（越大质量越好，速度越慢） |\n| `--solver` | str | euler | 求解器，可选 euler / midpoint / rk4 |\n| `--lambd` | float | 0.75 | 去噪强度 0~1 |\n| `--tau` | float | 0.5 | 音质保留阈值 |\n\n---\n\n## 📂 项目结构"},{"language":"text","snippet":"---\n\n## 🎯 使用场景推荐\n\n| 场景 | 推荐模式 | 建议参数 |\n|------|----------|----------|\n| 会议录音、电话录音、播客降噪 | VoiceFixer | `mode=1` + `--cuda` |\n| 老旧磁带、历史录音修复 | VoiceFixer | `mode=1` |\n| 音乐、演唱、人声提升高频细节 | AudioSR | `--hifi --model_name speech` |\n| 需要极致音质（音乐制作） | AudioSR | `--hifi --ddim_steps 100 --guidance_scale 4.0` |\n| 批量处理大量文件 | 任一模式 | 配合 `-o` 指定输出目录 |\n\n---\n\n## ⚠️ 注意事项\n\n1. **首次运行** VoiceFixer 或 AudioSR 时会自动下载模型，请保持网络畅通。\n2. AudioSR 处理速度较慢（尤其是 `ddim_steps` 较大时），建议使用 GPU。\n3. VoiceFixer 默认输出为 `.wav` 格式，AudioSR 默认输出为 48kHz `.wav`。\n4. 建议为不同任务分别创建输出目录，避免混淆。\n\n---\n\n## 🛠️ 未来计划（可选）\n\n- 支持更多音频增强模型\n- 添加图形界面（Gradio / Streamlit）\n- 支持批量配置文件\n- 集成到 OpenClaw 主框架\n- 支持实时音频流处理\n\n---\n\n## 📄 License\n\nApache License\n\n---\n\n## ❤️ 致谢\n\n- [AudioSR](https://github.com/haoheliu/versatile_audio_super_resolution.git) - 音频超级分辨率\n- [VoiceFixer](https://github.com/haoheliu/voicefixer.git) - 语音修复工具\n\n---\n\n**欢迎 Star ⭐ 支持项目！**\n\n如有问题或建议，欢迎提交 Issue 或 Pull Request。"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\r\nname: audio-enhancement-engine\r\ndescription: |\r\n  当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。\r\n  集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。\r\n  默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。\r\n  支持 wav、mp3、flac、m4a、ogg 等常见格式，完全本地运行，输出统一为高质量 WAV 文件。\r\n\r\n  【重要约束】仅处理音频文件或音频文件夹，其他文件（如视频、图片、文档、纯文本）一律不触发此技能。\r\n\r\n  常见触发口语（越多越好）：\r\n  - “帮我增强这个音频”\r\n  - “修复这个录音的音质”\r\n  - “给这个语音降噪”\r\n  - “把这个音频提升到高保真”\r\n  - “音乐音质增强 这个.mp3”\r\n  - “批量处理音频文件夹”\r\n  - “清理会议录音”\r\n  - “提升音频采样率到48kHz”\r\n  - “语音修复 这个 wav 文件”\r\n  - “高保真增强音频”\r\n  - “老旧录音修复”\r\n  - “音频增强 目录路径”\r\nmetadata:\r\n  openclaw:\r\n    requires:\r\n      bins:\r\n        - python\r\n    user-invocable: true\r\n---\r\n\r\n# Audio Enhancement Skill\r\n\r\n**功能**：本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\r\n\r\n### 触发时机（Triggers）\r\n- 用户提供音频文件（.wav、.mp3、.flac、.m4a、.ogg 等）或音频文件夹路径，并表达增强音质、修复、降噪、高保真等意图。\r\n- 用户说“音频增强”“修复录音”“降噪”“提升音质”“高保真”“48kHz”等关键词。\r\n- 支持单个文件处理或整个文件夹批量处理（支持递归子目录）。\r\n\r\n### 支持的两种增强模式\r\n1. **VoiceFixer 通用语音修复**（默认模式）\r\n   - 擅长语音降噪、提升清晰度、修复轻微失真。\r\n   - 推荐用于：会议录音、访谈、播客、语音笔记、老旧录音。\r\n\r\n2. **AudioSR 高保真音频超级分辨率**（启用 `--hifi` 时）\r\n   - 将音频提升至 48kHz，显著增加高频细节和整体保真度。\r\n   - 推荐用于：音乐、演唱、人声、需要高音质的场景。\r\n\r\n## 参数提取指南\r\n当决定调用此技能时，请从用户消息中准确提取以下参数：\r\n\r\n1. **`<输入路径>`** (必填): 用户提供的音频文件路径或文件夹路径（支持相对/绝对路径）。\r\n2. **`<输出路径>`** (选填): 用户指定的输出文件或目录路径。若未指定，默认在输入同级目录自动添加 `_enhanced` 后缀。\r\n3. **`<模式选择>`** (选填):\r\n   - 默认使用 VoiceFixer。\r\n   - 若用户提到“高保真”“音乐”“48kHz”“超分辨率”等，自动添加 `--hifi` 并使用 AudioSR。\r\n4. **VoiceFixer 专用参数**（默认模式）:\r\n   - `--mode`：0/1/2（推荐 1，默认 1）\r\n   - `--cuda`：是否使用 GPU\r\n   - `-r, --recursive`：是否递归子目录\r\n5. **AudioSR 专用参数**（`--hifi` 模式）:\r\n   - `--model_name`：`basic` 或 `speech`（人声推荐 speech）\r\n   - `--ddim_steps`：扩散步数（默认 50，建议 50-100）\r\n   - `--guidance_scale`：引导尺度（默认 3.5）\r\n   - `--seed`：随机种子（默认 42）\r\n   - `--device`：`cuda` 或 `cpu`\r\n\r\n### 执行步骤\r\n1. **解析路径**：识别用户提供的音频文件或文件夹路径。\r\n2. **模式判断**：根据用户意图判断使用 VoiceFixer（默认）还是 AudioSR（含 `--hifi`）。\r\n3. **默认目标**：若未指定输出路径，默认在输入目录生成带 `_enhanced_48k`（AudioSR）或 `_enhanced`（VoiceFixer）后缀的文件。\r\n4. **调用命令**：使用以下兼容性命令启动脚本（优先 `python3`，失败则 `python`）。脚本会自动检查环境、初始化模型并处理。\r\n\r\n   ```bash\r\n   (python3 scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>] [--re] [--nfe <数值>] [--solver <euler|midpoint|rk4>] [--lambd <数值>] [--tau <数值>]) || (python scripts/enhancer.py -i \"<输入路径>\" [-o \"<输出目录>\"] [-m <0|1|2>] [--cuda] [-r] [--hifi] [--model_name <basic|speech>] [--ddim_steps <数值>] [--guidance_scale <数值>] [--seed <数值>] [--device <cuda|cpu>] [--re] [--nfe <数值>] [--solver <euler|midpoint|rk4>] [--lambd <数值>] [--tau <数值>])"},{"path":"README.md","content":"### audio-enhancement-engine\n\n[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![Python](https://img.shields.io/badge/Python-3.8%2B-green.svg)](https://www.python.org/)\n[![OpenClaw](https://img.shields.io/badge/OpenClaw-Skill-orange.svg)](https://github.com/wangminrui2022)\n\n**audio-enhancement-engine** 本地音频增强与修复统一工具，集成 VoiceFixer（语音降噪/修复）和 AudioSR（高保真超级分辨率）。支持单文件与目录批量处理，自动适配最合适的增强模式，输出清晰、高质量的 48kHz WAV 文件。\n\n---\n\n```markdown\n# OpenClaw Audio Skill\n\n**一个简单高效的音频增强与修复命令行工具**\n\n支持两种主流音频增强技术：\n- **AudioSR**：高保真音频超级分辨率（将音频提升至 48kHz，增加细节与高频）\n- **VoiceFixer**：通用语音修复（降噪、提升清晰度、修复失真）\n\n---\n\n## ✨ 功能特点\n\n- 支持**单个文件**和**整个目录**批量处理\n- 同时集成两种专业音频增强模型\n- 通过 `--hifi` 一键切换高保真模式与语音修复模式\n- 支持 GPU 加速（CUDA）\n- 自动创建输出目录，智能添加 `_enhanced` 后缀\n- 详细的中文日志提示，处理状态清晰\n- 兼容性强（包含 NumPy 兼容补丁）\n- 支持递归处理子目录\n\n---\n\n## 📥 安装与使用\n\n### 1. 克隆项目\n\n```bash\ngit clone https://github.com/wangminrui2022/audio-enhancement-engine.git\n```\n\n> 首次运行 VoiceFixer 时会自动下载模型，AudioSR 同样会根据需要下载对应模型。\n\n### 2. 基本使用\n\n#### 方式一：默认使用 VoiceFixer（语音修复）\n\n```bash\n# 修复单个音频文件\npython scripts/enhancer.py -i input/audio.mp3\n\n# 修复整个目录\npython scripts/enhancer.py -i recordings/\n\n# 使用 GPU + 递归处理子目录\npython scripts/enhancer.py -i recordings/ --cuda -r\n```\n\n#### 方式二：使用 AudioSR 高保真增强（48kHz 超分辨率）\n\n```bash\n# 高保真增强单个文件\npython scripts/enhancer.py -i low_quality.wav --hifi\n\n# 高保真增强目录，并指定输出目录\npython scripts/enhancer.py -i music_folder/ -o enhanced_music/ --hifi\n\n# 使用更高参数提升质量\npython scripts/enhancer.py -i input.wav --hifi --ddim_steps 100 --guidance_scale 4.0\n```\n\n#### 方式三：Resemble-Enhance 模型构建的专业级语音增强工具\n\n```bash\n# 高保真增强单个文件\npython scripts/enhancer.py -i low_quality.wav --re\n\n# 高保真增强目录，并指定输出目录\npython scripts/enhancer.py -i music_folder/ -o enhanced_music/ --re\n\n# 使用更高参数提升质量\npython scripts/enhancer.py -i input.wav --re --nfe 64 --solver \"euler\" --lambd 0.75 --tau 0.5\n```\n\n\n---\n### **在 OpenClaw 聊天中**\n\n你可以直接对你的 Agent 说：\n\n帮我增强这个音频 recording.mp3”\n\n复这个会议录音，音质太差了”\n\n给这个音乐提升高保真音质 music.wav”\n\n把这个音频提升到48kHz”\n\n批量处理这个音频文件夹 audio_files/”\n\n高保真增强这个播客音频”\n\n帮我降噪这个语音笔记 voice.m4a”\n\n老旧录音修复 folder_path/”\n\n音乐音质增强 这个.mp3 --hifi”\n\n清理这个失真录音并提升清晰度”\n\n## 📋 命令行参数说明\n\n### 通用参数\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--input` | `-i` | str | **必填** | 输入文件或目录路径 |\n| `--output` | `-o` | str | None | 输出路径（文件或目录） |\n\n### VoiceFixer 参数（默认模式）\n\n| 参数 | 缩写 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `--mode` | `-m` | int | 1 | 增强模式 (0/1/2)，推荐使用 1 |\n| `--cuda` | | bool | False | 是否使用 GPU 加速 |\n| `--recursive` | `-r` | bool | False | 是否递归处理子目录 |\n\n### AudioSR 高保真参数（需添加 `--hifi`）\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `--hifi` | bool | False | 启用 AudioSR 高保真模式 |\n| `--model_name` | str | basic | 模型名称，可选 `basic` / `speech` |\n| `--ddim_steps` | int | 50 | 扩散步数（越大质量越好，速度越慢） |\n| `--guidance_scale` | float | 3.5 | 引导尺度 |\n| `--seed` | int | 42 | 随机种子 |\n| `--device` | str | None | 指定设备 `cuda` 或 `cpu` |\n\n### Resemble-Enhance 模型构建的专业级语音增强工具\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|-"},{"path":"scripts/resemble-enhance-0.0.1/README.md","content":"# Resemble Enhance\n\n[![PyPI](https://img.shields.io/pypi/v/resemble-enhance.svg)](https://pypi.org/project/resemble-enhance/)\n[![Hugging Face Space](https://img.shields.io/badge/Hugging%20Face%20%F0%9F%A4%97-Space-yellow)](https://huggingface.co/spaces/ResembleAI/resemble-enhance)\n[![License](https://img.shields.io/github/license/resemble-ai/Resemble-Enhance.svg)](https://github.com/resemble-ai/resemble-enhance/blob/main/LICENSE)\n\nhttps://github.com/resemble-ai/resemble-enhance/assets/660224/bc3ec943-e795-4646-b119-cce327c810f1\n\nResemble Enhance is an AI-powered tool that aims to improve the overall quality of speech by performing denoising and enhancement. It consists of two modules: a denoiser, which separates speech from a noisy audio, and an enhancer, which further boosts the perceptual audio quality by restoring audio distortions and extending the audio bandwidth. The two models are trained on high-quality 44.1kHz speech data that guarantees the enhancement of your speech with high quality.\n\n## Usage\n\n### Installation\n\n```bash\npip install resemble-enhance\n```\n\n### Enhance\n\n```\nresemble_enhance in_dir out_dir\n```\n\n### Denoise only\n\n```\nresemble_enhance in_dir out_dir --denoise_only\n```\n\n### Web Demo\n\nWe provide a web demo built with Gradio, you can try it out [here](https://huggingface.co/spaces/ResembleAI/resemble-enhance), or also run it locally:\n\n```\npython app.py\n```\n\n## Train your own model\n\n### Data Preparation\n\nYou need to prepare a foreground speech dataset and a background non-speech dataset. In addition, you need to prepare a RIR dataset ([examples](https://github.com/RoyJames/room-impulse-responses)).\n\n```bash\ndata\n├── fg\n│   ├── 00001.wav\n│   └── ...\n├── bg\n│   ├── 00001.wav\n│   └── ...\n└── rir\n    ├── 00001.npy\n    └── ...\n```\n\n### Training\n\n#### Denoiser Warmup\n\nThough the denoiser is trained jointly with the enhancer, it is recommended for a warmup training first.\n\n```bash\npython -m resemble_enhance.denoiser.train --yaml config/denoiser.yaml\n```\n\n#### Enhancer\n\nThen, you can train the enhancer in two stages. The first stage is to train the autoencoder and vocoder. And the second stage is to train the latent conditional flow matching (CFM) model.\n\n##### Stage 1\n\n```bash\npython -m resemble_enhance.enhancer.train --yaml config/enhancer_stage1.yaml\n```\n\n##### Stage 2\n\n```bash\npython -m resemble_enhance.enhancer.train --yaml config/enhancer_stage2.yaml\n```"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn77473086ppakmtqf0e4myp8981mhjd\",\n  \"slug\": \"audio-enhancement-engine\",\n  \"version\": \"1.0.5\",\n  \"publishedAt\": 1783087887476\n}"},{"path":"scripts/resemble-enhance-0.0.1/config/denoiser.yaml","content":"batch_size_per_gpu: 32\ntraining_seconds: 3.0"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。 集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。 默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。 支持 wav、mp3、flac、m4a、ogg 等常见 Skill: audio-enhancement-engine Owner: wangminrui2022 Summary: 当用户想要**音频增强**、**提升音质**、**修复录音**、**降噪**、**语音修复**、**高保真音频**、**48kHz超分辨率**、**清理会议录音**、**音乐音质提升**、**批量处理音频**时自动触发。 集成 **VoiceFixer**（通用语音修复）与 **AudioSR**（高保真音频超级分辨率到48kHz）两种专业技术，支持单个音频文件或整个目录批量处理。 默认使用 VoiceFixer 进行降噪和清晰度提升；当用户提到“高保真”“音乐增强”“提升采样率”“48kHz”等需求时，自动切换到 AudioSR 模式。 支持 wav、mp3、flac、m4a、ogg 等常见 Tags: latest:1.0.5 Version history: v1.0.5 | 2026-07-03T","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":924,"uniquenessScore":54,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T15:39:59.603Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T15:39:59.603Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T17:44:13.255Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}