agentCLAWHUBUnverified

Edge TTS — Text-to-Speech

Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English. Skill: Edge TTS — Text-to-Speech Owner: vincentlau2046-sudo Summary: Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English. Tags: latest:1.2.0 Version history: v1.2.0 | 2026-05-23T11:51:32.506Z | user Add --srt flag for native SRT subtitle generation from TTS timeline (exact text match, no ASR needed) v1.1.0 | 2026-05-23T08:5

OpenClaw

Rank

62

Safety

84

Downloads

1.1k

Updated

Oct 11, 2026

Version

1.2.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.1K downloadsadoption · observed Oct 11, 2026
Latest release
1.2.0release · observed May 23, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: medium.

clawhub skill install s177qxxtwn13trxmgd54c7a6wh83h7rc:tts-cosyvoice
  1. Setup complexity is MEDIUM. Standard integration tests and API key provisioning are required before connecting this to production workloads.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/snapshot"

Documentation

CLAWHUB

17,785 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: tts-cosyvoice
version: 1.0.0
description: "Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English."
metadata: { "openclaw": { "emoji": "🔊", "requires": { "bins": ["python3"] } } }
tags: ["tts", "text-to-speech", "edge-tts", "voice", "audio"]
---

# Edge TTS — Text-to-Speech

Use to convert text to speech using Microsoft Edge's free TTS service. No API key required.

## Voices

### Chinese (Recommended)
| Voice | Gender | Style |
|-------|--------|-------|
| `zh-CN-XiaoxiaoNeural` | Female | Warm, natural (default) |
| `zh-CN-YunxiNeural` | Male | Young, casual |
| `zh-CN-YunjianNeural` | Male | Professional, news |
| `zh-CN-XiaoyiNeural` | Female | Cute, young |
| `zh-CN-YunyangNeural` | Male | News broadcaster |
| `zh-CN-XiaochenNeural` | Female | Child |
| `zh-TW-HsiaoChenNeural` | Female | Traditional Chinese |

### English
| Voice | Gender | Style |
|-------|--------|-------|
| `en-US-JennyNeural` | Female | Friendly (default) |
| `en-US-GuyNeural` | Male | Professional |
| `en-GB-SoniaNeural` | Female | British |

## Script

```bash
{baseDir}/scripts/tts.py --text "Hello world" --output /tmp/output.mp3
```

### Options

| Option | Default | Description |
|--------|---------|-------------|
| `--text` | (required) | Text to speak |
| `--voice` | zh-CN-XiaoxiaoNeural | Voice ID |
| `--output` | /tmp/tts_output.mp3 | Output file path |
| `--rate` | 0% | Speed adjustment (-50% to +100%) |
| `--pitch` | 0Hz | Pitch adjustment (-50Hz to +50Hz) |
| `--volume` | 0% | Volume adjustment (-100% to +100%) |
| `--file` | | Read text from file instead of --text |

### Examples

```bash
# Basic Chinese TTS
{baseDir}/scripts/tts.py --text "你好,我是Nova" --output /tmp/hello.mp3

# Male voice, faster speech
{baseDir}/scripts/tts.py --text "Hello world" --voice en-US-GuyNeural --rate +20% --output /tmp/fast.mp3

# Read from file
{baseDir}/scripts/tts.py --file script.txt --output /tmp/script.mp3
```

## Available Voices

List all voices:
```bash
{baseDir}/scripts/tts.py --list-voices zh
{baseDir}/scripts/tts.py --list-voices en
{baseDir}/scripts/tts.py --list-voices ja
```

## Dependencies

- `edge-tts` (pip install)
- No API key needed — uses Microsoft Edge's free TTS service
- Requires internet connection

README.md

# TTS/ASR 语音模型部署指南

## 当前部署状态

| 组件 | 方案 | 状态 | 说明 |
|------|------|------|------|
| **TTS** | Edge TTS | ✅ 已部署 | 微软免费云端 TTS,无需 API Key |
| **ASR** | Whisper | ⏳ 待装 | pip install openai-whisper 依赖较多,国内网络慢 |

## TTS — Edge TTS

### 优势
- **零依赖**:只需 `edge-tts` Python 包(已装好)
- **免费**:无需 API Key,使用 Microsoft Edge 内置 TTS
- **高质量**:Neural 级别语音,中文自然度顶级
- **多语言**:50+ 语言,100+ 音色
- **轻量**:纯 Python,不占显存

### 中文推荐音色
| 音色 | 性别 | 风格 |
|------|------|------|
| `zh-CN-XiaoxiaoNeural` | 女 | 温暖自然(默认) |
| `zh-CN-YunxiNeural` | 男 | 年轻 casual |
| `zh-CN-YunyangNeural` | 男 | 新闻播报 |
| `zh-CN-XiaoyiNeural` | 女 | 可爱年轻 |

### 使用示例
```bash
~/.openclaw/plugin-skills/tts-cosyvoice/scripts/tts.py \
  --text "你好,我是Nova" \
  --output /tmp/hello.mp3
```

### 备选:CosyVoice(本地 TTS)
如果需要完全离线的 TTS:
```bash
pip install cosyvoice
# 从 ModelScope 下载模型
python -c "
from modelscope.hub.snapshot_download import snapshot_download
snapshot_download('iic/CosyVoice-300M-SFT', local_dir='./cosyvoice-models/sft')
"
```

## ASR — Whisper

### 当前状态
- `pip install openai-whisper` 依赖较多(torch_complex, kaldiio, librosa 等)
- 国内网络下载较慢
- 模型权重首次使用自动下载(base ~150MB)

### 安装(网络恢复后执行)
```bash
pip install openai-whisper
# 测试
python -c "import whisper; model = whisper.load_model('base'); print('OK')"
```

### 使用示例
```bash
~/.openclaw/plugin-skills/asr-funasr/scripts/asr.py \
  --input meeting.mp3 \
  --language zh \
  --model base
```

### 备选:FunASR SenseVoice(更快更准的中文 ASR)
如果 Whisper 安装困难:
```bash
pip install funasr modelscope
# 从 ModelScope 下载模型
python -c "
from funasr import AutoModel
model = AutoModel(model='iic/SenseVoiceSmall', trust_remote_code=True, device='cuda:0')
result = model.generate(input='test.wav')
print(result)
"
```

## 显存预算

| 场景 | ComfyUI | TTS | ASR | 总计 |
|------|---------|-----|-----|------|
| 图片生成 | 14GB | 0 | 0 | 14GB ✅ |
| 视频生成 | 13GB | 0 | 0 | 13GB ✅ |
| TTS | 0 | 0(云端) | 0 | 0 ✅ |
| ASR | 0 | 0 | 1-2GB | 1-2GB ✅ |
| TTS+ASR+ComfyUI | 13GB | 0 | 2GB | 15GB ✅ |

**Edge TTS 不占显存**(云端 API),ASR 仅 1-2GB(base 模型)。

## 与 keynote-video 集成

keynote-video 技能流程:
```
PPT → 内容理解(LLM)→ 讲稿生成 → TTS 语音合成 → 视频合成 → 音画合并
                                          ↑ Edge TTS
```

## OpenClaw 技能

| 技能 | 位置 | 功能 |
|------|------|------|
| tts-cosyvoice | ~/.openclaw/plugin-skills/tts-cosyvoice/ | 文字转语音 |
| asr-funasr | ~/.openclaw/plugin-skills/asr-funasr/ | 语音转文字 |

_meta.json

{
  "ownerId": "kn7dbzxmk4fh39kk5rbgsdj5nn82qjcs",
  "slug": "tts-cosyvoice",
  "version": "1.2.0",
  "publishedAt": 1779537092506
}

skill-card.md

## Description:

Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English.

This skill is ready for commercial/non-commercial use.

## Publisher:

[vincentlau2046-sudo](https://clawhub.ai/user/vincentlau2046-sudo)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and agents use this skill to convert text or text files into speech through Microsoft Edge TTS voices, with options for voice, rate, pitch, volume, language-filtered voice listing, and optional SRT subtitle generation.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill can send submitted text to Microsoft-operated cloud TTS infrastructure.

Mitigation: Do not submit sensitive text unless that data transfer is approved; use an offline TTS path for confidential content.

Risk: The release includes code-loading and setup patterns that could execute unreviewed third-party Python code.

Mitigation: Review before installing, use a dedicated virtual environment with pinned dependencies, remove the hard-coded /home/vincent sys.path override, and avoid trust_remote_code examples unless pinned and sandboxed.

## Reference(s):

- [ClawHub skill page](https://clawhub.ai/vincentlau2046-sudo/skills/tts-cosyvoice)

## Skill Output:

**Output Type(s):** [Guidance, Shell commands, Files]

**Output Format:** [Markdown guidance with bash commands; generated outputs are MP3 audio files and optional SRT subtitle files.]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Requires python3, the edge-tts Python package, and an internet connection to reach Microsoft-operated TTS infrastructure.]

## Skill Version(s):

1.2.0 (source: server release evidence; artifact frontmatter says 1.0.0)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 2d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/vincentlau2046-sudo/skills/tts-cosyvoice",
      "sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/skills/tts-cosyvoice",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T09:18:31.458Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T09:18:31.458Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.1K downloads",
      "href": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
      "sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T09:18:31.458Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.2.0",
      "href": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
      "sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-23T11:51:32.506Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.2.0",
      "description": "Add --srt flag for native SRT subtitle generation from TTS timeline (exact text match, no ASR needed)",
      "href": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
      "sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-23T11:51:32.506Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to Edge TTS — Text-to-Speech and adjacent AI workflows.