talkies
Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.
Rank
62
Safety
84
Downloads
1.9k
Updated
Oct 9, 2026
Version
1.3.18
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.9K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 1.9K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.3.18release · observed Aug 27, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:talkies- Install using `clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:talkies` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/psyb0t/talkies before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/snapshot"
Documentation
CLAWHUB
150,969 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: talkies
description: Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.
homepage: https://github.com/psyb0t/docker-talkies
user-invocable: true
permissions:
network: "Outbound HTTP to the configured TALKIES_URL; the Talkies server also fetches URLs supplied as file_path."
shell: "Documented setup and workflow examples invoke local curl, ffmpeg, and docker commands."
filesystem: "Reads and writes server-side staged files through /v1/files; the skill itself does not access the local filesystem."
metadata:
{ "openclaw": { "emoji": "🎙️", "primaryEnv": "TALKIES_URL", "requires": { "bins": ["docker", "curl"] } } }
---
# talkies
Self-hosted speech service — ASR and TTS, one container. OpenAI-compatible wire shape on both endpoints; point an OpenAI client at it, change the model slug, done.
ASR (`POST /v1/audio/transcriptions`): fourteen bundled slugs — `whisper-large-v3`, `whisper-large-v3-turbo`, `parakeet-tdt-0.6b-v3`, `nemotron-3.5-asr-0.6b`, `canary-180m-flash`, `canary-1b-flash`, `canary-qwen-2.5b`, four selectable English Sherpa Zipformer variants, `vosk-small-en-us-0.15`, and two phoneme recognizers, `wav2vec2-xlsr-53-espeak` and `zipa-ipa`, that return the IPA phones spoken instead of words.
TTS (`POST /v1/audio/speech`): 3 engines / 4 backends across 8 slugs — `kokoro-82m` (PyTorch) and `kokoro-82m-nvidia` (ONNX/ORT) with 41 baked voices across en/es/fr/hi/it/pt, plus the CUDA-only Qwen3-TTS family: `qwen3-tts-0.6b` / `qwen3-tts-1.7b` (voice cloning from reference clips), `qwen3-tts-0.6b-custom` / `qwen3-tts-1.7b-custom` (9 preset speakers), `qwen3-tts-1.7b-design` (voice from an NL description), plus the CUDA-only `chatterbox-turbo` (English only; 19 inline emotion tags; clones from a reference `.wav` with no transcript). Discover voices via `GET /v1/audio/voices`.
Extras: live PCM ASR over WebSocket, stereo diarization on transcription, URL `file_path` fetching, server-side file staging, MCP endpoint with 6 ASR-side tools, optional bearer-token auth.
For installation, configuration, and container setup, see [references/setup.md](references/setup.md).
## Security & safety
This skill is **not** low-risk to orchestrate blindly — it issues local shell commands (`curl`, `ffmpeg`, `docker` in the typical deployment/setup path) and outbound HTTP requests to whatever `$TALKIES_URL` points at:
- **Outbound HTTP to an operator-chosen host —_meta.json
{
"ownerId": "kn79dhvmpjng4rp2jjk8k0v5xx80ccbk",
"slug": "talkies",
"version": "1.3.18",
"publishedAt": 1787871031680
}references/setup.md
# talkies setup
## Requirements
- Docker
- `linux/amd64` host (no arm64 images — `nemo_toolkit[asr]` + chain doesn't resolve cleanly on aarch64)
- Optional: NVIDIA GPU + NVIDIA Container Toolkit for the CUDA image (required for `qwen3-tts-0.6b` voice cloning)
- ~3 GB disk for the CPU image, ~11 GB for the CUDA image
- Additional disk for selected model weights; set `TALKIES_ENABLED_MODELS` to avoid downloading the full registry
- ~4 GB RAM minimum (whisper-large-v3 needs the working set + overhead); 12 GB+ VRAM for the GPU-only models
## Quick Install
### CPU
Serves 2× Whisper + `canary-180m-flash` + `nemotron-3.5-asr-0.6b` (CPU-optimized, via parakeet.cpp), four selectable Sherpa Zipformer variants, and `vosk-small-en-us-0.15` for ASR, plus `kokoro-82m` and `kokoro-82m-nvidia` for TTS. The CUDA-only ASR models aren't worth running on CPU, and the Qwen3-TTS family is CUDA-only.
```bash
docker run -d --name talkies \
-v $HOME/talkies-data:/data \
-p 8000:8000 \
psyb0t/talkies:latest
```
### CUDA
Serves all fourteen ASR models plus all three TTS engines / 4 backends (`kokoro-82m`, `kokoro-82m-nvidia`, the 5 Qwen3-TTS slugs, and `chatterbox-turbo`). Requires the NVIDIA Container Toolkit on the host.
Nemotron runs through the SHA-256-pinned upstream parakeet.cpp v0.5.0 CUDA 12
bundle in this image, so both file transcription and native WebSocket sessions
use GPU offload. The matching CUDA 12.9 runtime libraries stay isolated under
`/opt/parakeet` from the image's Python ML stack.
```bash
docker run -d --name talkies \
--gpus all \
-v $HOME/talkies-data:/data \
-p 8000:8000 \
psyb0t/talkies:latest-cuda
```
The CUDA image expects `--gpus all`. Without a GPU assignment it retains its
`TALKIES_DEVICE=cuda` image default, so model loading fails rather than silently
falling back. To use its CPU-compatible subset for debugging, explicitly set
`-e TALKIES_DEVICE=cpu` and restrict `TALKIES_ENABLED_MODELS` to CPU-compatible
slugs; GPU-only Qwen3-TTS slugs remain unavailable.
**Verify:** `curl http://localhost:8000/healthz` returns `{"ok": true, "device": "...", "models": [...]}` once boot's done.
**First boot:** the entrypoint downloads every enabled model into `/data/models/<slug>/` and creates `/data/files/` + `/data/custom-voices/`. Bind-mount `/data` so subsequent restarts are no-ops. Restrict the download set with `TALKIES_ENABLED_MODELS` to avoid pulling everything.
## CPU vs CUDA Images
| Image | Tag | Platforms | Models served | Image size |
|---|---|---|---|---|
| CPU | `psyb0t/talkies:latest` | `linux/amd64` | 2× Whisper, Canary-180m-Flash, Nemotron-3.5-ASR, Sherpa Zipformer ×4, Vosk, wav2vec2 + ZIPA phoneme, Kokoro-82M ×2 runtimes | ~3 GB |
| CUDA | `psyb0t/talkies:latest-cuda` | `linux/amd64` | all fourteen ASR + Kokoro-82M ×2 runtimes + Qwen3-TTS ×5 + Chatterbox Turbo | ~11 GB |
The CPU image only ships ASR models that actually finish in a sane time without a GPU. Parakeet-TDT is autoregressive (slow on CPU). Canary-1skill-card.md
## Description: talkies helps agents use a self-hosted OpenAI-compatible speech service for transcription, live ASR, speech synthesis, subtitles, server-side file staging, and MCP-based ASR workflows. This skill is ready for commercial/non-commercial use. ## Publisher: [psyb0t](https://clawhub.ai/user/psyb0t) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agents use talkies to run speech-to-text, text-to-speech, subtitle generation, live transcription, and batch transcription workflows against a trusted self-hosted Talkies server. It is useful when an OpenAI-compatible audio API, optional MCP ASR tools, and local control over speech models are required. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Speech input, synthesized text, and voice-cloning reference samples are sent to the configured Talkies server. Mitigation: Point TALKIES_URL only at an instance the operator runs or explicitly trusts, prefer HTTPS outside localhost or a protected LAN, and avoid sending sensitive media or text to untrusted servers. Risk: An unconfigured Talkies deployment can expose the API, staged files, and management endpoints to anyone who can reach the port. Mitigation: Bind the service to localhost or a protected network, set TALKIES_AUTH_TOKEN for shared or public deployments, and use a reverse proxy with TLS and rate limiting when exposing it externally. Risk: Server-side URL file_path fetching can access operator-side network resources and persist downloaded media on the server. Mitigation: Enable TALKIES_BLOCK_PRIVATE_DOWNLOADS before accepting untrusted callers, pass only trusted URLs, and delete cached downloads after the workflow is complete. Risk: Server-side staged files are persistent, shared, and enumerable by callers with API access. Mitigation: Do not stage sensitive media on shared instances without authentication, only read or delete files created for the current workflow, and clean up staged files promptly. Risk: Voice cloning and expressive TTS can be misused to impersonate real people. Mitigation: Clone or synthesize a real person's voice only with explicit authorization and consent, and remove custom voice samples when no longer needed. Risk: Setup and workflow examples execute local docker, curl, and ffmpeg commands, and latest container tags may change over time. Mitigation: Review commands before execution, run them in an appropriate environment, and prefer pinned image digests for repeatable deployments. Risk: Debug logging can expose request and response content, including transcripts, TTS text, and voice reference transcripts. Mitigation: Keep production deployments at info or higher log levels and use debug logging only for local troubleshooting with synthetic or disposable data. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/psyb0t/skills/talkies) - [talkies setup guide](references/setup.md) - [Talkies repository](https://github.com/psyb0t/docker-talkies
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/psyb0t/skills/talkies",
"sourceUrl": "https://clawhub.ai/psyb0t/skills/talkies",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T22:59:11.240Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T22:59:11.240Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.9K downloads",
"href": "https://clawhub.ai/psyb0t/talkies",
"sourceUrl": "https://clawhub.ai/psyb0t/talkies",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T22:59:11.240Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.3.18",
"href": "https://clawhub.ai/psyb0t/talkies",
"sourceUrl": "https://clawhub.ai/psyb0t/talkies",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-27T22:50:31.680Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.3.18",
"description": "- Added two new phoneme recognizer models (wav2vec2-xlsr-53-espeak and ZIPA-IPA) to ASR, bringing the total bundled ASR slugs to 14. - Updated documentation to reflect the new phoneme recognition capabilities, which emit IPA phones. - Removed the skill-card.md file. - Clarified model lists and endpoints in overview and use-case sections.",
"href": "https://clawhub.ai/psyb0t/talkies",
"sourceUrl": "https://clawhub.ai/psyb0t/talkies",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-27T22:50:31.680Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
