kesha-voice-kit
Local multilingual voice toolkit — speech-to-text (STT), text-to-speech (TTS), speaker diarization, and language detection, over a CLI or an MCP server. Runs entirely offline on Apple Silicon, Linux, and Windows. No API keys, no cloud. NVIDIA Parakeet TDT for STT across 25 European languages, Kokoro-82M + Vosk-TTS for TTS in 9 languages, plus macOS AVSpeechSynthesizer for ~180 system voices with zero install. Skill: kesha-voice-kit Owner: drakulavich Summary: Local multilingual voice toolkit — speech-to-text (STT), text-to-speech (TTS), speaker diarization, and language detection, over a CLI or an MCP server. Runs entirely offline on Apple Silicon, Linux, and Windows. No API keys, no cloud. NVIDIA Parakeet TDT for STT across 25 European languages, Kokoro-82M + Vosk-TTS for TTS in 9 languages, plus macOS AVSpeechSynthesize
Rank
62
Safety
84
Downloads
1.6k
Updated
Oct 10, 2026
Version
1.6.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.6K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.6.1release · observed Aug 5, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17b3c9zks1vdxe5e7vt4f316h8571qe:kesha-voice-kit- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-drakulavich-kesha-voice-kit/snapshot"
Documentation
CLAWHUB
147,861 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: kesha-voice-kit
description: Local multilingual voice toolkit — speech-to-text (STT), text-to-speech (TTS), speaker diarization, and language detection, over a CLI or an MCP server. Runs entirely offline on Apple Silicon, Linux, and Windows. No API keys, no cloud. NVIDIA Parakeet TDT for STT across 25 European languages, Kokoro-82M + Vosk-TTS for TTS in 9 languages, plus macOS AVSpeechSynthesizer for ~180 system voices with zero install.
emoji: 🎙️
requires:
bins: [kesha]
install:
- kind: bash
cmd: bun add -g "@drakulavich/kesha-voice-kit"
- kind: bash
cmd: kesha install
---
# kesha-voice-kit
Local voice toolkit: transcribe voice messages to text, synthesize speech, detect language of audio or text. Fully offline after `kesha install`. No API keys, no per-minute billing.
**Trigger keywords for when to use this skill:** voice message, voice memo, voice note, .ogg, .opus, .wav, .mp3, audio file, transcribe, transcription, speech-to-text, STT, text-to-speech, TTS, synthesize speech, say, telegram voice note, whatsapp voice note, ogg-opus, opus, multilingual voice, multilingual ASR, language detection, speaker diarization, who said what, meeting transcript, MCP server, offline voice, privacy, Apple Silicon, CoreML.
## When to use
- **Voice memo arrived** (Telegram, WhatsApp, Slack, Signal .ogg/.opus/.m4a): transcribe with `kesha --json <path>` and branch on the detected language.
- **Need to send a voice note (Telegram, WhatsApp, Signal, Discord)**: synthesize directly into messenger-native OGG/Opus with `kesha say --format ogg-opus --out reply.ogg "<text>"`. Default is mono 24 kHz @ 32 kbps - what Telegram `sendVoice` expects. No WAV redirect and no `ffmpeg` round-trip.
- **Need local file playback/debug output**: WAV is still available with `kesha say --out reply.wav "<text>"`, but do not use WAV for Telegram voice replies. Auto-routes by detected language (Kokoro-82M for English, Vosk-TTS for Russian). On darwin-arm64, English Kokoro uses FluidAudio CoreML instead of ONNX. For other languages and ~180 more voices use `--voice macos-*` on macOS (zero model download).
- **Need to detect what language a file is in** before choosing a pipeline: `kesha --json audio.ogg` returns both audio-based and text-based language detection with confidence scores.
- **Need to capture your own voice** for transcription or as a voice-note source: `kesha record --out clip.wav` records up to 120s (override with `--max-seconds`) of mono 16 kHz WAV from the default microphone. Pipe straight into `kesha --json clip.wav` to close the loop.
## OpenClaw plugin setup
Install the plugin, then explicitly route OpenClaw audio understanding through the CLI model entry. The plugin registration makes Kesha discoverable, but real voice-message transcription uses `tools.media.audio.models` with a `type: "cli"` entry.
```bash
bun add -g @drakulavich/kesha-voice-kit
kesha install
openclaw plugins install @drakulavich/kesha-voice-kit
openclaw config patREADME.md
<p align="center">
<img src="https://github.com/drakulavich/kesha-voice-kit/raw/main/docs/assets/logo.png" alt="Kesha Voice Kit" width="200">
</p>
<h1 align="center">Kesha Voice Kit</h1>
<p align="center">
<a href="https://flakiness.io/Laputa/kesha-voice-kit"><img src="https://img.shields.io/endpoint?url=https%3A%2F%2Fflakiness.io%2Fapi%2Fbadge%3Finput%3D%257B%2522badgeToken%2522%253A%2522badge-2IKMRRqUxh9P3w8Ym3Szf0%2522%257D" alt="Tests"></a>
<a href="https://www.npmjs.com/package/@drakulavich/kesha-voice-kit"><img src="https://img.shields.io/npm/v/@drakulavich/kesha-voice-kit" alt="npm version"></a>
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-blue.svg" alt="License: MIT"></a>
<a href="https://bun.sh"><img src="https://img.shields.io/badge/runtime-Bun-f9f1e1?logo=bun" alt="Bun"></a>
</p>
<p align="center"><b>Give your local tools and LLM agents a voice.</b><br>Fast speech-to-text, text-to-speech, voice-activity detection, and language detection in one local-first CLI: Apple Silicon CoreML first, ONNX fallback on supported Linux/Windows builds.</p>
- **Transcribe locally** — [25 languages](docs/languages.md#speech-to-text-25), up to ~19x faster than Whisper on Apple Silicon, ~2.5x on CPU
- **Speak back** — text-to-speech in [9 languages](docs/languages.md#text-to-speech)
- **Plug into agents** — ship voice workflows as CLI commands, an MCP server, an <a href="docs/openclaw.md">OpenClaw</a> skill, or a <a href="docs/hermes.md">Hermes</a> agent
- **Small Rust engine** — single ~60MB binary, no ffmpeg, no Python, no native Node addons
<p align="center">
<img src="https://github.com/drakulavich/kesha-voice-kit/raw/main/demo.gif" alt="kesha demo — English + Russian transcription with automatic language detection" width="800">
</p>
## Quick Start
Runtime: **[Bun](https://bun.sh)** >= 1.3.0 · Platforms: macOS arm64, Linux x64, Windows x64. Linux and Windows run the ONNX engine — everything except microphone capture (`kesha record`), macOS system voices, speaker diarization, and text language detection, which need Apple frameworks.
```bash
# 1. Install Bun (skip if you have it) — Linux & macOS:
curl -fsSL https://bun.sh/install | bash # or: brew install oven-sh/bun/bun
# Windows: powershell -c "irm bun.sh/install.ps1 | iex"
# if `bun --version` fails, reload PATH: exec $SHELL -l
# 2. Install Kesha:
bun add -g @drakulavich/kesha-voice-kit
kesha --version # confirms `kesha` resolved on PATH
kesha install --plan # preview exact download/disk sizes first — downloads nothing
kesha install # ~2.5 GB on Linux/Windows; ~0.6 GB on Apple Silicon, whose CoreML
# engine uses a different, smaller model set. Explicit — never automatic.
# No progress bar during the model step; can take several minutes.
# Prefer a guided wizard? `kesha init` walks through the _meta.json
{
"ownerId": "kn70xsptbaknapzrxhhsqepa4x80ynkp",
"slug": "kesha-voice-kit",
"version": "1.6.1",
"publishedAt": 1785948561169
}skill-card.md
## Description: Kesha Voice Kit lets agents transcribe audio, synthesize speech, perform speaker diarization, and detect language locally through a CLI or MCP server. This skill is ready for commercial/non-commercial use. ## Publisher: [drakulavich](https://clawhub.ai/user/drakulavich) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agent builders use this skill to add local voice-message transcription, text-to-speech voice-note generation, language detection, and MCP or OpenClaw audio workflows without cloud speech APIs. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The install path uses mutable remote and global installers with broad local code execution authority. Mitigation: Review install commands before execution, prefer package-manager or pinned installation methods, avoid piping remote installers directly into a shell, and run the tool as a normal non-admin user. Risk: Installation persists a global CLI plus model assets and can change OpenClaw audio or TTS behavior. Mitigation: Preview planned downloads with installation plan commands, verify OpenClaw audio and TTS configuration after install, and keep the CLI route explicit in user configuration. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/drakulavich/skills/kesha-voice-kit) - [npm package](https://www.npmjs.com/package/@drakulavich/kesha-voice-kit) - [Bun runtime](https://bun.sh) ## Skill Output: **Output Type(s):** [text, markdown, JSON, shell commands, configuration, guidance] **Output Format:** [Markdown with inline shell commands and JSON configuration examples] **Output Parameters:** [1D] **Other Properties Related to Output:** [Agent-facing guidance may direct local CLI calls that produce transcripts, timestamped JSON, language metadata, or generated audio files.] ## Skill Version(s): 1.6.1 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/drakulavich/skills/kesha-voice-kit",
"sourceUrl": "https://clawhub.ai/drakulavich/skills/kesha-voice-kit",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T06:05:55.014Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-drakulavich-kesha-voice-kit/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-drakulavich-kesha-voice-kit/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T06:05:55.014Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.6K downloads",
"href": "https://clawhub.ai/drakulavich/kesha-voice-kit",
"sourceUrl": "https://clawhub.ai/drakulavich/kesha-voice-kit",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T06:05:55.014Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.6.1",
"href": "https://clawhub.ai/drakulavich/kesha-voice-kit",
"sourceUrl": "https://clawhub.ai/drakulavich/kesha-voice-kit",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-05T16:49:21.169Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-drakulavich-kesha-voice-kit/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-drakulavich-kesha-voice-kit/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.6.1",
"description": "Retag 1.6.0 content onto every topic tag (they were still pinned to 1.5.0, so category browsing served the old description) and restore the Kesha Voice Kit display name. No content change from 1.6.0.",
"href": "https://clawhub.ai/drakulavich/kesha-voice-kit",
"sourceUrl": "https://clawhub.ai/drakulavich/kesha-voice-kit",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-05T16:49:21.169Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
