Edge TTS — Text-to-Speech
Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English. Skill: Edge TTS — Text-to-Speech Owner: vincentlau2046-sudo Summary: Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English. Tags: latest:1.2.0 Version history: v1.2.0 | 2026-05-23T11:51:32.506Z | user Add --srt flag for native SRT subtitle generation from TTS timeline (exact text match, no ASR needed) v1.1.0 | 2026-05-23T08:5
Rank
62
Safety
84
Downloads
1.1k
Updated
Oct 11, 2026
Version
1.2.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.2.0release · observed May 23, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: medium.
clawhub skill install s177qxxtwn13trxmgd54c7a6wh83h7rc:tts-cosyvoice- Setup complexity is MEDIUM. Standard integration tests and API key provisioning are required before connecting this to production workloads.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/snapshot"
Documentation
CLAWHUB
17,785 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: tts-cosyvoice
version: 1.0.0
description: "Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English."
metadata: { "openclaw": { "emoji": "🔊", "requires": { "bins": ["python3"] } } }
tags: ["tts", "text-to-speech", "edge-tts", "voice", "audio"]
---
# Edge TTS — Text-to-Speech
Use to convert text to speech using Microsoft Edge's free TTS service. No API key required.
## Voices
### Chinese (Recommended)
| Voice | Gender | Style |
|-------|--------|-------|
| `zh-CN-XiaoxiaoNeural` | Female | Warm, natural (default) |
| `zh-CN-YunxiNeural` | Male | Young, casual |
| `zh-CN-YunjianNeural` | Male | Professional, news |
| `zh-CN-XiaoyiNeural` | Female | Cute, young |
| `zh-CN-YunyangNeural` | Male | News broadcaster |
| `zh-CN-XiaochenNeural` | Female | Child |
| `zh-TW-HsiaoChenNeural` | Female | Traditional Chinese |
### English
| Voice | Gender | Style |
|-------|--------|-------|
| `en-US-JennyNeural` | Female | Friendly (default) |
| `en-US-GuyNeural` | Male | Professional |
| `en-GB-SoniaNeural` | Female | British |
## Script
```bash
{baseDir}/scripts/tts.py --text "Hello world" --output /tmp/output.mp3
```
### Options
| Option | Default | Description |
|--------|---------|-------------|
| `--text` | (required) | Text to speak |
| `--voice` | zh-CN-XiaoxiaoNeural | Voice ID |
| `--output` | /tmp/tts_output.mp3 | Output file path |
| `--rate` | 0% | Speed adjustment (-50% to +100%) |
| `--pitch` | 0Hz | Pitch adjustment (-50Hz to +50Hz) |
| `--volume` | 0% | Volume adjustment (-100% to +100%) |
| `--file` | | Read text from file instead of --text |
### Examples
```bash
# Basic Chinese TTS
{baseDir}/scripts/tts.py --text "你好,我是Nova" --output /tmp/hello.mp3
# Male voice, faster speech
{baseDir}/scripts/tts.py --text "Hello world" --voice en-US-GuyNeural --rate +20% --output /tmp/fast.mp3
# Read from file
{baseDir}/scripts/tts.py --file script.txt --output /tmp/script.mp3
```
## Available Voices
List all voices:
```bash
{baseDir}/scripts/tts.py --list-voices zh
{baseDir}/scripts/tts.py --list-voices en
{baseDir}/scripts/tts.py --list-voices ja
```
## Dependencies
- `edge-tts` (pip install)
- No API key needed — uses Microsoft Edge's free TTS service
- Requires internet connectionREADME.md
# TTS/ASR 语音模型部署指南
## 当前部署状态
| 组件 | 方案 | 状态 | 说明 |
|------|------|------|------|
| **TTS** | Edge TTS | ✅ 已部署 | 微软免费云端 TTS,无需 API Key |
| **ASR** | Whisper | ⏳ 待装 | pip install openai-whisper 依赖较多,国内网络慢 |
## TTS — Edge TTS
### 优势
- **零依赖**:只需 `edge-tts` Python 包(已装好)
- **免费**:无需 API Key,使用 Microsoft Edge 内置 TTS
- **高质量**:Neural 级别语音,中文自然度顶级
- **多语言**:50+ 语言,100+ 音色
- **轻量**:纯 Python,不占显存
### 中文推荐音色
| 音色 | 性别 | 风格 |
|------|------|------|
| `zh-CN-XiaoxiaoNeural` | 女 | 温暖自然(默认) |
| `zh-CN-YunxiNeural` | 男 | 年轻 casual |
| `zh-CN-YunyangNeural` | 男 | 新闻播报 |
| `zh-CN-XiaoyiNeural` | 女 | 可爱年轻 |
### 使用示例
```bash
~/.openclaw/plugin-skills/tts-cosyvoice/scripts/tts.py \
--text "你好,我是Nova" \
--output /tmp/hello.mp3
```
### 备选:CosyVoice(本地 TTS)
如果需要完全离线的 TTS:
```bash
pip install cosyvoice
# 从 ModelScope 下载模型
python -c "
from modelscope.hub.snapshot_download import snapshot_download
snapshot_download('iic/CosyVoice-300M-SFT', local_dir='./cosyvoice-models/sft')
"
```
## ASR — Whisper
### 当前状态
- `pip install openai-whisper` 依赖较多(torch_complex, kaldiio, librosa 等)
- 国内网络下载较慢
- 模型权重首次使用自动下载(base ~150MB)
### 安装(网络恢复后执行)
```bash
pip install openai-whisper
# 测试
python -c "import whisper; model = whisper.load_model('base'); print('OK')"
```
### 使用示例
```bash
~/.openclaw/plugin-skills/asr-funasr/scripts/asr.py \
--input meeting.mp3 \
--language zh \
--model base
```
### 备选:FunASR SenseVoice(更快更准的中文 ASR)
如果 Whisper 安装困难:
```bash
pip install funasr modelscope
# 从 ModelScope 下载模型
python -c "
from funasr import AutoModel
model = AutoModel(model='iic/SenseVoiceSmall', trust_remote_code=True, device='cuda:0')
result = model.generate(input='test.wav')
print(result)
"
```
## 显存预算
| 场景 | ComfyUI | TTS | ASR | 总计 |
|------|---------|-----|-----|------|
| 图片生成 | 14GB | 0 | 0 | 14GB ✅ |
| 视频生成 | 13GB | 0 | 0 | 13GB ✅ |
| TTS | 0 | 0(云端) | 0 | 0 ✅ |
| ASR | 0 | 0 | 1-2GB | 1-2GB ✅ |
| TTS+ASR+ComfyUI | 13GB | 0 | 2GB | 15GB ✅ |
**Edge TTS 不占显存**(云端 API),ASR 仅 1-2GB(base 模型)。
## 与 keynote-video 集成
keynote-video 技能流程:
```
PPT → 内容理解(LLM)→ 讲稿生成 → TTS 语音合成 → 视频合成 → 音画合并
↑ Edge TTS
```
## OpenClaw 技能
| 技能 | 位置 | 功能 |
|------|------|------|
| tts-cosyvoice | ~/.openclaw/plugin-skills/tts-cosyvoice/ | 文字转语音 |
| asr-funasr | ~/.openclaw/plugin-skills/asr-funasr/ | 语音转文字 |_meta.json
{
"ownerId": "kn7dbzxmk4fh39kk5rbgsdj5nn82qjcs",
"slug": "tts-cosyvoice",
"version": "1.2.0",
"publishedAt": 1779537092506
}skill-card.md
## Description: Text-to-Speech via Edge TTS (Microsoft Azure voices). Free, no API key needed, supports 100+ voices in 50+ languages including Chinese and English. This skill is ready for commercial/non-commercial use. ## Publisher: [vincentlau2046-sudo](https://clawhub.ai/user/vincentlau2046-sudo) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agents use this skill to convert text or text files into speech through Microsoft Edge TTS voices, with options for voice, rate, pitch, volume, language-filtered voice listing, and optional SRT subtitle generation. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The skill can send submitted text to Microsoft-operated cloud TTS infrastructure. Mitigation: Do not submit sensitive text unless that data transfer is approved; use an offline TTS path for confidential content. Risk: The release includes code-loading and setup patterns that could execute unreviewed third-party Python code. Mitigation: Review before installing, use a dedicated virtual environment with pinned dependencies, remove the hard-coded /home/vincent sys.path override, and avoid trust_remote_code examples unless pinned and sandboxed. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/vincentlau2046-sudo/skills/tts-cosyvoice) ## Skill Output: **Output Type(s):** [Guidance, Shell commands, Files] **Output Format:** [Markdown guidance with bash commands; generated outputs are MP3 audio files and optional SRT subtitle files.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Requires python3, the edge-tts Python package, and an internet connection to reach Microsoft-operated TTS infrastructure.] ## Skill Version(s): 1.2.0 (source: server release evidence; artifact frontmatter says 1.0.0) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/vincentlau2046-sudo/skills/tts-cosyvoice",
"sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/skills/tts-cosyvoice",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T09:18:31.458Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T09:18:31.458Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.1K downloads",
"href": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
"sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T09:18:31.458Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.2.0",
"href": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
"sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-23T11:51:32.506Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-vincentlau2046-sudo-tts-cosyvoice/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.2.0",
"description": "Add --srt flag for native SRT subtitle generation from TTS timeline (exact text match, no ASR needed)",
"href": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
"sourceUrl": "https://clawhub.ai/vincentlau2046-sudo/tts-cosyvoice",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-23T11:51:32.506Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
