Zvukogram
Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio.... Skill: Zvukogram Owner: erview Summary: Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio.... Tags: latest:1.1.6 Version history: v1.1.6 | 2026-03-29T16:59:42.275Z | user Sync latest production learnings into the public skill: explicitly document no-silent-truncation for >1000-char /text limits, strengthen SSML
Rank
62
Safety
84
Downloads
1.8k
Updated
Oct 10, 2026
Version
1.1.6
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.8K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.8K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.1.6release · observed Mar 29, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s175h0jdx5g1z9bpzgdxrfjban83t7pq:zvukogram- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/snapshot"
Documentation
CLAWHUB
148,647 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: zvukogram
description: Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio. Supports speed control, stress marks, English word transcription, audio fragment merging, rich SSML references, and podcast-oriented TTS patterns.
metadata:
openclaw:
requires:
env: [ZVUKOGRAM_TOKEN, ZVUKOGRAM_EMAIL]
credentials: [zvukogram_api]
---
# Zvukogram TTS
Speech generation via Zvukogram API with SSML markup support.
## Requirements
To use this skill, you need:
- **Zvukogram API token** — get it at https://zvukogram.com/
- **Zvukogram account email**
### Setup
Create file `~/.config/zvukogram/config.json`:
```bash
mkdir -p ~/.config/zvukogram
```
```json
{
"token": "your_api_token_here",
"email": "[email protected]"
}
```
Or use environment variables:
```bash
export ZVUKOGRAM_TOKEN=your_api_token_here
export [email protected]
```
## Quick Start
```bash
# Simple TTS
python3 scripts/tts.py --text "Hello, world!" --voice Алена --output hello.mp3
# With +20% speed
python3 scripts/tts.py --text "Fast text" --voice Алена --speed 1.2 --output fast.mp3
# Check balance
python3 scripts/balance.py
```
## Features
- **TTS generation** — text to speech
- **SSML support** — stress marks, pauses, speed
- **Audio merging** — combine fragments via ffmpeg
- **Transcription** — proper pronunciation of English words
## SSML Markup
### Stress Marks
Use `+` before stressed vowel:
```
З+амок — stress on "a"
зам+ок — stress on "o"
```
### Aliases (Transcription)
```xml
<sub alias="Оупен Эй Ай">OpenAI</sub>
<sub alias="Самсунг">Samsung</sub>
<sub alias="Ал+ьтман">Альтман</sub>
```
### Speed
```xml
<prosody rate="1.2">20% faster</prosody>
<prosody rate="fast">Fast text</prosody>
```
### Pauses
```xml
<break time="500ms"/>
```
## Available Voices
- **Алена** — female, neutral (recommended)
- **Андрей** — male, neutral (recommended)
- **Александра** — female, soft
- **Антон** — male, business
Full list: see [references/VOICES.md](references/VOICES.md)
## Examples
See [references/EXAMPLES.md](references/EXAMPLES.md) for:
- Dialogs and podcasts
- News voiceover
- Voice notifications
- Long texts
## Transcription
See [references/TRANSCRIPTION.md](references/TRANSCRIPTION.md) for proper pronunciation:
- OpenAI → Оупен Эй Ай
- GPT → Джи Пи Ти
- Samsung → Самсунг
- Altman → Ал+ьтман
## SSML Reference
- Full, agent-readable reference (recommended): [references/SSML.md](references/SSML.md)
- `say-as` modes with extra patterns: [references/say-as.md](references/say-as.md)
- Pronunciation & transcription patterns (`+`, `<sub>`, `<phoneme>`): [references/pronunciation-patterns.md](references/pronunciation-patterns.md)
- Podcast-oriented SSML patterns: [references/podcast-examples.md](references/podcast-examples.md)
- Quick lookup: [references/SSML_CHEATSHEET.md](references/SSML_CHEATSHEET.md)
- Official Zvukogram _meta.json
{
"ownerId": "kn7asmvryqycathxsgmqa6ht6581f5eq",
"slug": "zvukogram",
"version": "1.1.6",
"publishedAt": 1774803582275
}references/API.md
# Zvukogram API — agent reference Source (official): https://zvukogram.com/node/api/ This doc is the **practical contract** for calling Zvukogram HTTP API. It focuses on what an agent needs: which endpoint to pick, key parameters, limits, and response fields. Base URL: - `https://zvukogram.com/index.php?r=api` Endpoints used most: - `POST/GET .../text` - `POST/GET .../longtext` - `POST/GET .../subs` - `POST/GET .../result` - `GET .../voices` - `POST/GET .../balance` - `POST/GET .../delete` ## 1) Authentication (required everywhere) Every request must include: - `token` — your API key (from личный кабинет) - `email` — account email Transport formats supported (pick any): 1) `POST application/x-www-form-urlencoded` (most common) 2) `POST application/json` 3) `GET` query params ## 2) Which method to use: /text vs /longtext vs /subs ### `/text` (1 step) Use when: - You need **fast** results for **short text**. Hard limit: - `text` max **1000 characters** per request. Behavior: - Response is JSON with a direct `file` URL immediately. ### `/longtext` (2 steps) Use when: - Text is **long** (up to **1,000,000 characters**) OR you want server-side chunking/queue. Behavior: 1) Send to `/longtext` → get `id`, `status=0` 2) Poll `/result` with that `id` until `status=1` (ready) or `status=-1` (error) Polling guidance from docs: - every **2s** for short texts - every **5s** for very long texts (books) ### `/subs` (2 steps + subtitles/timecodes) Use when: - You need **timed cuts** / subtitles alignment for video/presentations. Works like `/longtext`, but `/result` additionally includes: - `cuts` — array of fragments (links) when using `/subs` (also mentioned with “obrezka” tag in docs) Extra params specific to `/subs`: - `speed_type` (int: 1 or 2) - `speed_floor` (int) If you don’t need timecodes, prefer `/text` or `/longtext`. ## 3) Common request parameters ### Required - `token` - `email` - `voice` — voice name, e.g. `Matthew plus`, `Мартын` (see `/voices`) - `text` — input text (plain text or SSML-inline text) ### Voice / expressiveness - `speed` (float, default `1`): **0.1 … 2.0** - 0.8 slower, 1.3 faster - `pitch` (int, default `0`): **-20 … 20** - `style` (string): voice style (not supported by all voices) - docs mention examples like `newscast`, `cheerful`, `sad` - legacy name `emotion` still works - `styledegree` (string/number): intensity of the style; works only with `style` - `role` (string): voice role, e.g. `YoungAdultMale`, `OlderAdultFemale` (voice-dependent) ### Pauses and loudness (API-level controls) These are useful when you **don’t** want to insert SSML `<break>` everywhere. - `pause_sentence` (int ms): pause between sentences - `pause_paragraph` (int ms): pause between paragraphs - `volume` (int, default `100`): **10 … 200** - `effect` (string): audio effect (voice-dependent), example in docs: `car` ### Output format / quality - `format` (string, default `mp3`): `mp3`, `wav`, `ogg`, `opus`, `flac`
references/chunking-and-method-choice.md
# /text vs /longtext vs chunking — practical guide Source: https://zvukogram.com/node/api/ ## 1) Decision rule (fast & reliable) 1) If the text chunk is **<= 1000 chars** and you want the result **immediately** → use **`/text`**. 2) If total text is **up to 1,000,000 chars** and you can wait / poll → use **`/longtext`**. 3) If you need **timecodes / cuts** → use **`/subs`**. 4) If you want **instant results for a long text** (and control over pacing/voices) → **chunk yourself** and call `/text` multiple times. ## 2) Why chunking is often better for podcasts Chunking (many `/text` calls) gives you: - “Edit points” between segments (easy to cut/reorder) - Multi-voice flows (one voice per request) + merge - More stable prosody (each chunk is a clean sentence/paragraph) - No queue waiting for `/longtext` (you get each part as soon as it’s ready) Trade-off: - More API calls; you need merging and slightly more bookkeeping. ## 3) Safe chunking strategy Goals: - Keep each chunk **<= 1000 characters** (hard API limit for `/text`). - Avoid breaking inside SSML tags. - Prefer splitting on **paragraphs → sentences**. Recommended approach: 1) Normalize whitespace. 2) Split by blank lines (paragraph boundaries). 3) Within a paragraph, split by sentence punctuation (`.`, `!`, `?`, `…`). 4) Accumulate sentences until adding one would exceed your limit (use ~900–950 chars as a safety buffer). ### SSML-aware rule If you use SSML tags like `<sub>...</sub>` or `<say-as ...>...</say-as>`, never cut inside a tag pair. If you must chunk SSML-heavy text, chunk **before** tagging or ensure each chunk is well-formed XML fragments. ## 4) Merging audio Two common patterns: ### A) ffmpeg concat demuxer (most reliable) Create `list.txt`: ```text file 'part_001.mp3' file 'part_002.mp3' file 'part_003.mp3' ``` Then: ```bash ffmpeg -y -f concat -safe 0 -i list.txt -acodec copy out.mp3 ``` ### B) concat protocol (works for some mp3, less universal) ```bash ffmpeg -y -i "concat:part1.mp3|part2.mp3|part3.mp3" -acodec copy out.mp3 ``` ## 5) Polling /result correctly Only needed for `/longtext` or `/subs`. - Start polling `/result` with `id`. - Typical interval: **2 seconds**. - For very long jobs: **5 seconds**. - Stop when `status` is: - `1` → ready (`file` present) - `-1` → error (`error` present) ## 6) Non-negotiable production rule If text exceeds `/text` limits, do **not** silently truncate it. That is a production bug, not a fallback strategy. Use SSML-safe chunking or switch to `/longtext`. ## 6) Limits recap - `/text`: max **1000 characters** per request - `/longtext` and `/subs`: max **1,000,000 characters** (From official API docs.)
references/EXAMPLES.md
# Usage Examples
## Example 1: Simple Greeting
**Command:**
```bash
python3 skills/zvukogram/scripts/tts.py \
--text "Привет, Сергей! Как дела?" \
--voice Алена \
--output hello.mp3
```
## Example 2: News Voiceover
**Text (news.txt):**
```
Компания Оупен Эй Ай представила новую модель Джи Пи Ти 5.
Сэм Ал+ьтман заявил о прорыве в области искусственного интеллекта.
```
**Command:**
```bash
python3 skills/zvukogram/scripts/tts.py \
--file news.txt \
--voice Алена \
--speed 1.1 \
--output news.mp3
```
## Example 3: Dialog (Podcast)
**Dialog script (dialog.txt):**
```
[Алена|1.2] Доброе утро! Это подкаст Эй Ай Дейли.
[Андрей|1.2] Привет! Рад быть в эфире.
[Алена|1.2] Начнём с главных новостей.
[Андрей|1.2] В мире ИИ сегодня много интересного.
```
**Generation script:**
```bash
#!/bin/bash
TMP_DIR="/tmp/podcast_$(date +%s)"
mkdir -p $TMP_DIR
# Read dialog and generate fragments
while IFS= read -r line; do
if [[ $line =~ ^\[(.+)\|(.+)\]\ (.+)$ ]]; then
voice="${BASH_REMATCH[1]}"
speed="${BASH_REMATCH[2]}"
text="${BASH_REMATCH[3]}"
python3 skills/zvukogram/scripts/tts.py \
--text "$text" \
--voice "$voice" \
--speed "$speed" \
--output "$TMP_DIR/$(date +%s%N).mp3"
fi
done < dialog.txt
# Merge
python3 skills/zvukogram/scripts/merge.py \
$TMP_DIR/*.mp3 \
--output podcast.mp3
# Cleanup
rm -rf $TMP_DIR
```
## Example 4: Voice Notification
```bash
python3 skills/zvukogram/scripts/tts.py \
--text "Напоминание: встреча через 15 минут" \
--voice Алена \
--speed 1.1 \
--output reminder.mp3
```
## Example 5: Long Text
For texts longer than 1000 characters, use splitting:
```python
#!/usr/bin/env python3
import sys
sys.path.insert(0, 'skills/zvukogram/scripts')
from tts import generate_tts, download_audio, load_config
def split_text(text, max_len=900):
"""Split text by sentences"""
sentences = text.replace('! ', '!|').replace('? ', '?|').replace('. ', '.|').split('|')
parts = []
current = ""
for sent in sentences:
if len(current) + len(sent) < max_len:
current += sent + " "
else:
parts.append(current.strip())
current = sent + " "
if current:
parts.append(current.strip())
return parts
# Load text
with open('long_document.txt') as f:
text = f.read()
# Split and generate
config = load_config()
parts = split_text(text)
files = []
for i, part in enumerate(parts):
print(f"Генерация части {i+1}/{len(parts)}...")
url = generate_tts(part, "Алена", config["token"], config["email"], 1.0, "mp3")
if url:
filename = f"/tmp/part_{i:03d}.mp3"
download_audio(url, filename)
files.append(filename)
# Merge
import subprocess
with open('/tmp/merge_list.txt', 'w') as f:
for file in files:
f.write(f"file '{file}'\n")
subprocess.run([
'ffmpeg', '-y', '-f', 'concat', '-safe', '0',
'-i', '/tmp/merAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/erview/skills/zvukogram",
"sourceUrl": "https://clawhub.ai/erview/skills/zvukogram",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T02:57:31.342Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T02:57:31.342Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.8K downloads",
"href": "https://clawhub.ai/erview/zvukogram",
"sourceUrl": "https://clawhub.ai/erview/zvukogram",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T02:57:31.342Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.1.6",
"href": "https://clawhub.ai/erview/zvukogram",
"sourceUrl": "https://clawhub.ai/erview/zvukogram",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-03-29T16:59:42.275Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.1.6",
"description": "Sync latest production learnings into the public skill: explicitly document no-silent-truncation for >1000-char /text limits, strengthen SSML pipeline guidance to preserve useful inline tags when downstream runtimes support them, and keep wrapper-tag handling explicit. Includes prior enriched API/SSML/say-as/pronunciation/podcast references.",
"href": "https://clawhub.ai/erview/zvukogram",
"sourceUrl": "https://clawhub.ai/erview/zvukogram",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-03-29T16:59:42.275Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
