agentCLAWHUBUnverified

Zvukogram

Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio.... Skill: Zvukogram Owner: erview Summary: Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio.... Tags: latest:1.1.6 Version history: v1.1.6 | 2026-03-29T16:59:42.275Z | user Sync latest production learnings into the public skill: explicitly document no-silent-truncation for >1000-char /text limits, strengthen SSML

OpenClaw

Rank

62

Safety

84

Downloads

1.8k

Updated

Oct 10, 2026

Version

1.1.6

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.8K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.8K downloadsadoption · observed Oct 10, 2026
Latest release
1.1.6release · observed Mar 29, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s175h0jdx5g1z9bpzgdxrfjban83t7pq:zvukogram
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/snapshot"

Documentation

CLAWHUB

148,647 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: zvukogram
description: Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio. Supports speed control, stress marks, English word transcription, audio fragment merging, rich SSML references, and podcast-oriented TTS patterns.
metadata:
  openclaw:
    requires:
      env: [ZVUKOGRAM_TOKEN, ZVUKOGRAM_EMAIL]
      credentials: [zvukogram_api]
---

# Zvukogram TTS

Speech generation via Zvukogram API with SSML markup support.

## Requirements

To use this skill, you need:
- **Zvukogram API token** — get it at https://zvukogram.com/
- **Zvukogram account email**

### Setup

Create file `~/.config/zvukogram/config.json`:
```bash
mkdir -p ~/.config/zvukogram
```

```json
{
  "token": "your_api_token_here",
  "email": "[email protected]"
}
```

Or use environment variables:
```bash
export ZVUKOGRAM_TOKEN=your_api_token_here
export [email protected]
```

## Quick Start

```bash
# Simple TTS
python3 scripts/tts.py --text "Hello, world!" --voice Алена --output hello.mp3

# With +20% speed
python3 scripts/tts.py --text "Fast text" --voice Алена --speed 1.2 --output fast.mp3

# Check balance
python3 scripts/balance.py
```

## Features

- **TTS generation** — text to speech
- **SSML support** — stress marks, pauses, speed
- **Audio merging** — combine fragments via ffmpeg
- **Transcription** — proper pronunciation of English words

## SSML Markup

### Stress Marks
Use `+` before stressed vowel:
```
З+амок — stress on "a"
зам+ок — stress on "o"
```

### Aliases (Transcription)
```xml
<sub alias="Оупен Эй Ай">OpenAI</sub>
<sub alias="Самсунг">Samsung</sub>
<sub alias="Ал+ьтман">Альтман</sub>
```

### Speed
```xml
<prosody rate="1.2">20% faster</prosody>
<prosody rate="fast">Fast text</prosody>
```

### Pauses
```xml
<break time="500ms"/>
```

## Available Voices

- **Алена** — female, neutral (recommended)
- **Андрей** — male, neutral (recommended)
- **Александра** — female, soft
- **Антон** — male, business

Full list: see [references/VOICES.md](references/VOICES.md)

## Examples

See [references/EXAMPLES.md](references/EXAMPLES.md) for:
- Dialogs and podcasts
- News voiceover
- Voice notifications
- Long texts

## Transcription

See [references/TRANSCRIPTION.md](references/TRANSCRIPTION.md) for proper pronunciation:
- OpenAI → Оупен Эй Ай
- GPT → Джи Пи Ти
- Samsung → Самсунг
- Altman → Ал+ьтман

## SSML Reference

- Full, agent-readable reference (recommended): [references/SSML.md](references/SSML.md)
- `say-as` modes with extra patterns: [references/say-as.md](references/say-as.md)
- Pronunciation & transcription patterns (`+`, `<sub>`, `<phoneme>`): [references/pronunciation-patterns.md](references/pronunciation-patterns.md)
- Podcast-oriented SSML patterns: [references/podcast-examples.md](references/podcast-examples.md)
- Quick lookup: [references/SSML_CHEATSHEET.md](references/SSML_CHEATSHEET.md)
- Official Zvukogram 

_meta.json

{
  "ownerId": "kn7asmvryqycathxsgmqa6ht6581f5eq",
  "slug": "zvukogram",
  "version": "1.1.6",
  "publishedAt": 1774803582275
}

references/API.md

# Zvukogram API — agent reference

Source (official): https://zvukogram.com/node/api/

This doc is the **practical contract** for calling Zvukogram HTTP API. It focuses on what an agent needs: which endpoint to pick, key parameters, limits, and response fields.

Base URL:

- `https://zvukogram.com/index.php?r=api`

Endpoints used most:

- `POST/GET .../text`
- `POST/GET .../longtext`
- `POST/GET .../subs`
- `POST/GET .../result`
- `GET .../voices`
- `POST/GET .../balance`
- `POST/GET .../delete`

## 1) Authentication (required everywhere)

Every request must include:

- `token` — your API key (from личный кабинет)
- `email` — account email

Transport formats supported (pick any):

1) `POST application/x-www-form-urlencoded` (most common)
2) `POST application/json`
3) `GET` query params

## 2) Which method to use: /text vs /longtext vs /subs

### `/text` (1 step)

Use when:
- You need **fast** results for **short text**.

Hard limit:
- `text` max **1000 characters** per request.

Behavior:
- Response is JSON with a direct `file` URL immediately.

### `/longtext` (2 steps)

Use when:
- Text is **long** (up to **1,000,000 characters**) OR you want server-side chunking/queue.

Behavior:
1) Send to `/longtext` → get `id`, `status=0`
2) Poll `/result` with that `id` until `status=1` (ready) or `status=-1` (error)

Polling guidance from docs:
- every **2s** for short texts
- every **5s** for very long texts (books)

### `/subs` (2 steps + subtitles/timecodes)

Use when:
- You need **timed cuts** / subtitles alignment for video/presentations.

Works like `/longtext`, but `/result` additionally includes:
- `cuts` — array of fragments (links) when using `/subs` (also mentioned with “obrezka” tag in docs)

Extra params specific to `/subs`:
- `speed_type` (int: 1 or 2)
- `speed_floor` (int)

If you don’t need timecodes, prefer `/text` or `/longtext`.

## 3) Common request parameters

### Required

- `token`
- `email`
- `voice` — voice name, e.g. `Matthew plus`, `Мартын` (see `/voices`)
- `text` — input text (plain text or SSML-inline text)

### Voice / expressiveness

- `speed` (float, default `1`): **0.1 … 2.0**
  - 0.8 slower, 1.3 faster
- `pitch` (int, default `0`): **-20 … 20**
- `style` (string): voice style (not supported by all voices)
  - docs mention examples like `newscast`, `cheerful`, `sad`
  - legacy name `emotion` still works
- `styledegree` (string/number): intensity of the style; works only with `style`
- `role` (string): voice role, e.g. `YoungAdultMale`, `OlderAdultFemale` (voice-dependent)

### Pauses and loudness (API-level controls)

These are useful when you **don’t** want to insert SSML `<break>` everywhere.

- `pause_sentence` (int ms): pause between sentences
- `pause_paragraph` (int ms): pause between paragraphs
- `volume` (int, default `100`): **10 … 200**
- `effect` (string): audio effect (voice-dependent), example in docs: `car`

### Output format / quality

- `format` (string, default `mp3`): `mp3`, `wav`, `ogg`, `opus`, `flac`

references/chunking-and-method-choice.md

# /text vs /longtext vs chunking — practical guide

Source: https://zvukogram.com/node/api/

## 1) Decision rule (fast & reliable)

1) If the text chunk is **<= 1000 chars** and you want the result **immediately** → use **`/text`**.
2) If total text is **up to 1,000,000 chars** and you can wait / poll → use **`/longtext`**.
3) If you need **timecodes / cuts** → use **`/subs`**.
4) If you want **instant results for a long text** (and control over pacing/voices) → **chunk yourself** and call `/text` multiple times.

## 2) Why chunking is often better for podcasts

Chunking (many `/text` calls) gives you:

- “Edit points” between segments (easy to cut/reorder)
- Multi-voice flows (one voice per request) + merge
- More stable prosody (each chunk is a clean sentence/paragraph)
- No queue waiting for `/longtext` (you get each part as soon as it’s ready)

Trade-off:
- More API calls; you need merging and slightly more bookkeeping.

## 3) Safe chunking strategy

Goals:
- Keep each chunk **<= 1000 characters** (hard API limit for `/text`).
- Avoid breaking inside SSML tags.
- Prefer splitting on **paragraphs → sentences**.

Recommended approach:

1) Normalize whitespace.
2) Split by blank lines (paragraph boundaries).
3) Within a paragraph, split by sentence punctuation (`.`, `!`, `?`, `…`).
4) Accumulate sentences until adding one would exceed your limit (use ~900–950 chars as a safety buffer).

### SSML-aware rule

If you use SSML tags like `<sub>...</sub>` or `<say-as ...>...</say-as>`, never cut inside a tag pair.

If you must chunk SSML-heavy text, chunk **before** tagging or ensure each chunk is well-formed XML fragments.

## 4) Merging audio

Two common patterns:

### A) ffmpeg concat demuxer (most reliable)

Create `list.txt`:

```text
file 'part_001.mp3'
file 'part_002.mp3'
file 'part_003.mp3'
```

Then:

```bash
ffmpeg -y -f concat -safe 0 -i list.txt -acodec copy out.mp3
```

### B) concat protocol (works for some mp3, less universal)

```bash
ffmpeg -y -i "concat:part1.mp3|part2.mp3|part3.mp3" -acodec copy out.mp3
```

## 5) Polling /result correctly

Only needed for `/longtext` or `/subs`.

- Start polling `/result` with `id`.
- Typical interval: **2 seconds**.
- For very long jobs: **5 seconds**.
- Stop when `status` is:
  - `1` → ready (`file` present)
  - `-1` → error (`error` present)

## 6) Non-negotiable production rule

If text exceeds `/text` limits, do **not** silently truncate it.
That is a production bug, not a fallback strategy.
Use SSML-safe chunking or switch to `/longtext`.

## 6) Limits recap

- `/text`: max **1000 characters** per request
- `/longtext` and `/subs`: max **1,000,000 characters**

(From official API docs.)

references/EXAMPLES.md

# Usage Examples

## Example 1: Simple Greeting

**Command:**
```bash
python3 skills/zvukogram/scripts/tts.py \
  --text "Привет, Сергей! Как дела?" \
  --voice Алена \
  --output hello.mp3
```

## Example 2: News Voiceover

**Text (news.txt):**
```
Компания Оупен Эй Ай представила новую модель Джи Пи Ти 5.
Сэм Ал+ьтман заявил о прорыве в области искусственного интеллекта.
```

**Command:**
```bash
python3 skills/zvukogram/scripts/tts.py \
  --file news.txt \
  --voice Алена \
  --speed 1.1 \
  --output news.mp3
```

## Example 3: Dialog (Podcast)

**Dialog script (dialog.txt):**
```
[Алена|1.2] Доброе утро! Это подкаст Эй Ай Дейли.
[Андрей|1.2] Привет! Рад быть в эфире.
[Алена|1.2] Начнём с главных новостей.
[Андрей|1.2] В мире ИИ сегодня много интересного.
```

**Generation script:**
```bash
#!/bin/bash
TMP_DIR="/tmp/podcast_$(date +%s)"
mkdir -p $TMP_DIR

# Read dialog and generate fragments
while IFS= read -r line; do
    if [[ $line =~ ^\[(.+)\|(.+)\]\ (.+)$ ]]; then
        voice="${BASH_REMATCH[1]}"
        speed="${BASH_REMATCH[2]}"
        text="${BASH_REMATCH[3]}"
        
        python3 skills/zvukogram/scripts/tts.py \
          --text "$text" \
          --voice "$voice" \
          --speed "$speed" \
          --output "$TMP_DIR/$(date +%s%N).mp3"
    fi
done < dialog.txt

# Merge
python3 skills/zvukogram/scripts/merge.py \
  $TMP_DIR/*.mp3 \
  --output podcast.mp3

# Cleanup
rm -rf $TMP_DIR
```

## Example 4: Voice Notification

```bash
python3 skills/zvukogram/scripts/tts.py \
  --text "Напоминание: встреча через 15 минут" \
  --voice Алена \
  --speed 1.1 \
  --output reminder.mp3
```

## Example 5: Long Text

For texts longer than 1000 characters, use splitting:

```python
#!/usr/bin/env python3
import sys
sys.path.insert(0, 'skills/zvukogram/scripts')
from tts import generate_tts, download_audio, load_config

def split_text(text, max_len=900):
    """Split text by sentences"""
    sentences = text.replace('! ', '!|').replace('? ', '?|').replace('. ', '.|').split('|')
    parts = []
    current = ""
    
    for sent in sentences:
        if len(current) + len(sent) < max_len:
            current += sent + " "
        else:
            parts.append(current.strip())
            current = sent + " "
    
    if current:
        parts.append(current.strip())
    
    return parts

# Load text
with open('long_document.txt') as f:
    text = f.read()

# Split and generate
config = load_config()
parts = split_text(text)
files = []

for i, part in enumerate(parts):
    print(f"Генерация части {i+1}/{len(parts)}...")
    url = generate_tts(part, "Алена", config["token"], config["email"], 1.0, "mp3")
    if url:
        filename = f"/tmp/part_{i:03d}.mp3"
        download_audio(url, filename)
        files.append(filename)

# Merge
import subprocess
with open('/tmp/merge_list.txt', 'w') as f:
    for file in files:
        f.write(f"file '{file}'\n")

subprocess.run([
    'ffmpeg', '-y', '-f', 'concat', '-safe', '0',
    '-i', '/tmp/mer
Github ReposUpdated 9h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/erview/skills/zvukogram",
      "sourceUrl": "https://clawhub.ai/erview/skills/zvukogram",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T02:57:31.342Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T02:57:31.342Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.8K downloads",
      "href": "https://clawhub.ai/erview/zvukogram",
      "sourceUrl": "https://clawhub.ai/erview/zvukogram",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T02:57:31.342Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.1.6",
      "href": "https://clawhub.ai/erview/zvukogram",
      "sourceUrl": "https://clawhub.ai/erview/zvukogram",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-03-29T16:59:42.275Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-erview-zvukogram/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.1.6",
      "description": "Sync latest production learnings into the public skill: explicitly document no-silent-truncation for >1000-char /text limits, strengthen SSML pipeline guidance to preserve useful inline tags when downstream runtimes support them, and keep wrapper-tag handling explicit. Includes prior enriched API/SSML/say-as/pronunciation/podcast references.",
      "href": "https://clawhub.ai/erview/zvukogram",
      "sourceUrl": "https://clawhub.ai/erview/zvukogram",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-03-29T16:59:42.275Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to Zvukogram and adjacent AI workflows.