{"id":"6aabf4ae-dead-4b0a-be08-37f903acd722","entityType":"agent","slug":"clawhub-psyb0t-talkies","name":"talkies","canonicalUrl":"https://www.xpersona.co/agent/clawhub-psyb0t-talkies","canonicalPath":"/agent/clawhub-psyb0t-talkies","generatedAt":"2026-10-10T07:40:04.942Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T22:59:11.240Z","emptyReason":null},"description":"Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.9K downloads reported by the source. Last updated 10/9/2026.","installCommand":"clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:talkies","sourceUrl":"https://clawhub.ai/psyb0t/talkies","homepage":"https://clawhub.ai/psyb0t/skills/talkies","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/psyb0t/talkies","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/psyb0t/skills/talkies","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":66,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"talkies technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-09T22:59:11.240Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T22:59:11.240Z","emptyReason":null},"stars":null,"forks":null,"downloads":1919,"packageName":null,"latestVersion":"1.3.18","tractionLabel":"1.9K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T22:59:11.211Z","emptyReason":null},"lastUpdatedAt":"2026-10-09T22:59:11.240Z","lastCrawledAt":"2026-10-09T22:59:11.211Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-10T22:59:11.211Z","lastVerifiedAt":null,"highlights":[{"version":"1.3.18","createdAt":"2026-08-27T22:50:31.680Z","changelog":"- Added two new phoneme recognizer models (wav2vec2-xlsr-53-espeak and ZIPA-IPA) to ASR, bringing the total bundled ASR slugs to 14. - Updated documentation to reflect the new phoneme recognition capabilities, which emit IPA phones. - Removed the skill-card.md file. - Clarified model lists and endpoints in overview and use-case sections.","fileCount":5,"zipByteSize":32020},{"version":"1.3.17","createdAt":"2026-08-21T10:02:34.940Z","changelog":"talkies 1.3.17 - Removed skill-card.md, consolidating documentation into SKILL.md. - No functional changes; this update is documentation-only.","fileCount":5,"zipByteSize":31071},{"version":"1.3.16","createdAt":"2026-08-08T20:36:10.316Z","changelog":"- Updated documentation in SKILL.md and references/setup.md. - Removed obsolete skill-card.md file. - No changes to skill functionality; this is a documentation and cleanup update.","fileCount":5,"zipByteSize":31285},{"version":"1.3.15","createdAt":"2026-08-07T21:34:03.388Z","changelog":"- Updated documentation in references/setup.md. - Removed the skill-card.md file. - No changes to skill runtime or functionality; update is documentation-only.","fileCount":5,"zipByteSize":30806},{"version":"1.3.14","createdAt":"2026-08-07T20:02:45.130Z","changelog":"Chatterbox Turbo TTS and backend integration: - Added support for the CUDA-only \"Chatterbox Turbo\" TTS engine, offering English synthesis, 19 inline emotion tags, and transcript-free voice cloning. - Updated documentation and endpoints to reflect 3 TTS engines / 4 backends. - Listed new usage scenarios for Chatterbox Turbo, including emotive delivery with inline tags and cloning without needing transcripts. - Clarified engine limitations and compatibility in \"When To Use\" and \"When NOT To Use\" sections. - Removed legacy or obsolete documentation (skill-card.md).","fileCount":5,"zipByteSize":30692},{"version":"1.3.13","createdAt":"2026-08-06T03:27:11.966Z","changelog":"talkies 1.3.13 - Revised permissions and metadata formatting for consistency and clarity in SKILL.md. - Expanded SKILL.md feature summary and updated description of MCP endpoint (now referencing six ASR/file-staging MCP tools). - Removed skill-card.md, consolidating documentation. - Minor wording and formatting improvements in setup, safety, and quick start documentation.","fileCount":5,"zipByteSize":29725},{"version":"1.3.12","createdAt":"2026-08-06T02:40:18.124Z","changelog":"- Removed the skill-card.md file. - Updated references/setup.md with new or revised setup information. - Documentation cleanup, with all essential usage and setup details now consolidated in references/setup.md.","fileCount":5,"zipByteSize":29823},{"version":"1.3.11","createdAt":"2026-08-01T18:55:05.371Z","changelog":"- Removed the skill-card.md file. - Minor documentation updates in SKILL.md; no user-facing functionality changes documented.","fileCount":5,"zipByteSize":29374}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:talkies","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:talkies` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/psyb0t/talkies before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T07:40:04.938Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-talkies/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T22:59:11.240Z","emptyReason":null},"readme":"Skill: talkies\n\nOwner: psyb0t\n\nSummary: Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.\n\nTags: latest:1.3.18\n\nVersion history:\n\nv1.3.18 | 2026-08-27T22:50:31.680Z | auto\n\n- Added two new phoneme recognizer models (wav2vec2-xlsr-53-espeak and ZIPA-IPA) to ASR, bringing the total bundled ASR slugs to 14.\n- Updated documentation to reflect the new phoneme recognition capabilities, which emit IPA phones.\n- Removed the skill-card.md file.\n- Clarified model lists and endpoints in overview and use-case sections.\n\nv1.3.17 | 2026-08-21T10:02:34.940Z | auto\n\ntalkies 1.3.17\n\n- Removed skill-card.md, consolidating documentation into SKILL.md.\n- No functional changes; this update is documentation-only.\n\nv1.3.16 | 2026-08-08T20:36:10.316Z | auto\n\n- Updated documentation in SKILL.md and references/setup.md.\n- Removed obsolete skill-card.md file.\n- No changes to skill functionality; this is a documentation and cleanup update.\n\nv1.3.15 | 2026-08-07T21:34:03.388Z | auto\n\n- Updated documentation in references/setup.md.\n- Removed the skill-card.md file.\n- No changes to skill runtime or functionality; update is documentation-only.\n\nv1.3.14 | 2026-08-07T20:02:45.130Z | auto\n\nChatterbox Turbo TTS and backend integration:\n\n- Added support for the CUDA-only \"Chatterbox Turbo\" TTS engine, offering English synthesis, 19 inline emotion tags, and transcript-free voice cloning.\n- Updated documentation and endpoints to reflect 3 TTS engines / 4 backends.\n- Listed new usage scenarios for Chatterbox Turbo, including emotive delivery with inline tags and cloning without needing transcripts.\n- Clarified engine limitations and compatibility in \"When To Use\" and \"When NOT To Use\" sections.\n- Removed legacy or obsolete documentation (skill-card.md).\n\nv1.3.13 | 2026-08-06T03:27:11.966Z | auto\n\ntalkies 1.3.13\n\n- Revised permissions and metadata formatting for consistency and clarity in SKILL.md.\n- Expanded SKILL.md feature summary and updated description of MCP endpoint (now referencing six ASR/file-staging MCP tools).\n- Removed skill-card.md, consolidating documentation.\n- Minor wording and formatting improvements in setup, safety, and quick start documentation.\n\nv1.3.12 | 2026-08-06T02:40:18.124Z | auto\n\n- Removed the skill-card.md file.\n- Updated references/setup.md with new or revised setup information.\n- Documentation cleanup, with all essential usage and setup details now consolidated in references/setup.md.\n\nv1.3.11 | 2026-08-01T18:55:05.371Z | auto\n\n- Removed the skill-card.md file.\n- Minor documentation updates in SKILL.md; no user-facing functionality changes documented.\n\nv1.3.10 | 2026-07-31T09:45:02.885Z | auto\n\ntalkies 1.3.10\n\n- Documentation maintenance: SKILL.md and references/setup.md updated.\n- Removed skill-card.md file.\n- No functional or API changes.\n\nv1.3.9 | 2026-07-30T21:07:37.997Z | auto\n\ntalkies 1.3.9\n\n- Expanded ASR support from 7 to 12 bundled models, including Sherpa-ONNX and Vosk.\n- Updated documentation to reflect new ASR model options.\n- Removed the unused skill-card.md file for cleanup.\n- Minor updates in setup documentation.\n\nv1.3.8 | 2026-07-30T02:02:29.528Z | auto\n\ntalkies 1.3.8\n\n- Removed obsolete skill-card.md for streamlined documentation.\n- Updated SKILL.md and references/setup.md for improved clarity and accuracy.\n- No behavioral changes to skill endpoints or usage.\n\nv1.3.7 | 2026-07-30T01:08:21.148Z | auto\n\nVersion 1.3.7\n\n- Adds support for live PCM ASR over WebSocket via /v1/audio/transcriptions/stream, enabling streaming speech-to-text.\n- SKILL.md clarifies availability of real-time ASR endpoint and corrects endpoint documentation.\n- Updated references/setup.md to cover changes and usage related to live streaming.\n- Removes obsolete skill-card.md file.\n\nv1.3.6 | 2026-07-26T04:16:44.261Z | auto\n\nVersion 1.3.6\n\n- Internal documentation update: SKILL.md refined for clarity and detail.\n- Removed skill-card.md from the project. \n- No public-facing feature or behavioral changes to functionality.\n\nv1.3.5 | 2026-07-26T02:41:58.898Z | auto\n\ntalkies 1.3.5\n\n- Updated documentation in SKILL.md: clarified security and file staging notes, improved explanations of shared file handling and authentication defaults.\n- Removed skill-card.md file.\n\nv1.3.4 | 2026-07-26T01:46:38.683Z | auto\n\ntalkies 1.3.4\n\n- Expanded security & safety section to more strongly warn about data leaving your host, outgoing HTTP, file deletes, and voice synthesis consent.\n- Clarified that file deletion is destructive and not safe in multi-user/shared environments (no per-caller file ownership).\n- Emphasized that TTS `input` and cloning samples are always sent to remote server; recommend trusted deployments and HTTPS.\n- Removed the outdated skill-card.md file.\n\nv1.3.3 | 2026-07-25T23:33:23.101Z | auto\n\n# talkies 1.3.3\n\n- Added explicit `permissions` section to clarify security, HTTP, shell usage, and file operations.\n- Updated documentation to include important legal/ethical notice for Qwen3-TTS voice cloning; links to extended setup guidance.\n- Minor clarifications to server-side URL fetching and voice usage.\n- Removed the obsolete `skill-card.md` file.\n\nv1.3.2 | 2026-07-25T22:21:37.487Z | auto\n\n**Security warning and documentation changes.**\n\n- Added prominent security and safety disclaimer to documentation, detailing risks of unauthenticated access, shell command execution, and server-side file persistence.\n- Clarified that all API requests interact with an operator-specified server, and that remote servers are not vetted by the skill.\n- Documented that `file_path` URLs are fetched and cached server-side, not client-side.\n- Reduced the listed number of ASR backends from eight to seven in documentation (no functional code change indicated).\n- Removed obsolete skill-card documentation file.\n\nv1.3.1 | 2026-07-25T00:11:01.980Z | auto\n\ntalkies 1.3.1\n\n- Added support for Nemotron-3.5-ASR-0.6B (multilingual, CPU-optimized ASR).\n- Expanded TTS: new ONNX runtime via `kokoro-82m-nvidia`; Qwen3-TTS now exposes multiple slugs for different modes (preset speakers, voice design).\n- Updated documentation to reflect 8 ASR backends and revamped TTS backend matrix.\n- Removed legacy `skill-card.md` file.\n- General documentation refinements and improved setup details.\n\nv1.3.0 | 2026-05-28T18:10:21.362Z | auto\n\n- Removes the Distil-Whisper-Large-v3 automatic speech recognition (ASR) backend; now lists 6 ASR models instead of 7.\n- Updates documentation to reflect the new set of supported ASR models.\n- Keeps all other functionality, features, and usage the same.\n\nv1.2.0 | 2026-05-28T16:49:50.606Z | auto\n\ntalkies 1.2.0\n\n- Added support for Qwen3-TTS-0.6B voice cloning TTS engine (CUDA-only, works with user-provided `.wav` reference files in `/data/custom-voices/`).\n- `/v1/audio/speech` now supports two engines: Kokoro-82M (41 built-in voices) and Qwen3-TTS-0.6B (voice cloning).\n- `GET /v1/audio/voices` now discovers both Kokoro built-in and Qwen3-TTS custom voices.\n- Updated documentation and usage notes for voice cloning, new supported languages, and limitations (e.g. CUDA requirement for Qwen3-TTS).\n- Clarified OpenAI endpoint compatibility and handling of additional fields.\n\nv1.1.0 | 2026-05-28T14:18:23.058Z | auto\n\n**Talkies 1.1.0 — Adds TTS support (Kokoro-82M) alongside existing ASR**\n\n- Adds OpenAI-compatible text-to-speech endpoint (`/v1/audio/speech`) using the Kokoro-82M model with 41 voices across 6 languages.\n- New endpoint `/v1/audio/voices` for discovering supported TTS voices.\n- Updates description and documentation to reflect speech synthesis support; previous features remain (Whisper, Parakeet, Canary, diarization, URL fetch, etc.).\n- Maintains full OpenAI wire format compatibility for both transcription and TTS.\n- No API breaking changes for existing ASR users.\n\nv1.0.0 | 2026-05-28T11:41:22.409Z | auto\n\nInitial release of talkies — self-hosted, OpenAI-compatible speech-to-text API.\n\n- Compatible with OpenAI’s `/v1/audio/transcriptions` wire format; supports seven open ASR models (Whisper, Parakeet, Canary).\n- Features stereo diarization, server-side URL file fetching, multi-model backend, and optional bearer authentication.\n- Supports various audio formats and can generate plain transcripts, subtitles (SRT/VTT), or verbose JSON outputs with timestamps.\n- Drop-in replacement for OpenAI’s API; just change the endpoint URL and model slug.\n- See documentation for setup, supported models, and usage examples.\n\nArchive index:\n\nArchive v1.3.18: 5 files, 32020 bytes\n\nFiles: references/setup.md (24606b), scripts/bulk_transcribe.sh (3497b), skill-card.md (4012b), SKILL.md (49489b), _meta.json (127b)\n\nFile v1.3.18:SKILL.md\n\n---\nname: talkies\ndescription: Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.\nhomepage: https://github.com/psyb0t/docker-talkies\nuser-invocable: true\npermissions:\n  network: \"Outbound HTTP to the configured TALKIES_URL; the Talkies server also fetches URLs supplied as file_path.\"\n  shell: \"Documented setup and workflow examples invoke local curl, ffmpeg, and docker commands.\"\n  filesystem: \"Reads and writes server-side staged files through /v1/files; the skill itself does not access the local filesystem.\"\nmetadata:\n  { \"openclaw\": { \"emoji\": \"🎙️\", \"primaryEnv\": \"TALKIES_URL\", \"requires\": { \"bins\": [\"docker\", \"curl\"] } } }\n---\n\n# talkies\n\nSelf-hosted speech service — ASR and TTS, one container. OpenAI-compatible wire shape on both endpoints; point an OpenAI client at it, change the model slug, done.\n\nASR (`POST /v1/audio/transcriptions`): fourteen bundled slugs — `whisper-large-v3`, `whisper-large-v3-turbo`, `parakeet-tdt-0.6b-v3`, `nemotron-3.5-asr-0.6b`, `canary-180m-flash`, `canary-1b-flash`, `canary-qwen-2.5b`, four selectable English Sherpa Zipformer variants, `vosk-small-en-us-0.15`, and two phoneme recognizers, `wav2vec2-xlsr-53-espeak` and `zipa-ipa`, that return the IPA phones spoken instead of words.\n\nTTS (`POST /v1/audio/speech`): 3 engines / 4 backends across 8 slugs — `kokoro-82m` (PyTorch) and `kokoro-82m-nvidia` (ONNX/ORT) with 41 baked voices across en/es/fr/hi/it/pt, plus the CUDA-only Qwen3-TTS family: `qwen3-tts-0.6b` / `qwen3-tts-1.7b` (voice cloning from reference clips), `qwen3-tts-0.6b-custom` / `qwen3-tts-1.7b-custom` (9 preset speakers), `qwen3-tts-1.7b-design` (voice from an NL description), plus the CUDA-only `chatterbox-turbo` (English only; 19 inline emotion tags; clones from a reference `.wav` with no transcript). Discover voices via `GET /v1/audio/voices`.\n\nExtras: live PCM ASR over WebSocket, stereo diarization on transcription, URL `file_path` fetching, server-side file staging, MCP endpoint with 6 ASR-side tools, optional bearer-token auth.\n\nFor installation, configuration, and container setup, see [references/setup.md](references/setup.md).\n\n## Security & safety\n\nThis skill is **not** low-risk to orchestrate blindly — it issues local shell commands (`curl`, `ffmpeg`, `docker` in the typical deployment/setup path) and outbound HTTP requests to whatever `$TALKIES_URL` points at:\n\n- **Outbound HTTP to an operator-chosen host — data leaves your host.** Every command in this skill, including TTS `input` text and voice-cloning reference samples, is sent to whatever `$TALKIES_URL` points at — a server the skill does not control or vet. Point it only at an instance you run or explicitly trust; prefer HTTPS.\n- **`file_path` URL fetches happen server-side, not client-side.** Passing a URL causes the *talkies server* to download it — see [URL `file_path` (Download + Cache)](#url-file_path-download--cache) below.\n- **Staged files persist server-side until explicitly deleted**, and are enumerable by anyone who can reach the API — see [Server-Side File Staging](#server-side-file-staging-v1files) below.\n- **Shared file staging** — staged files (`/v1/files`) live in one namespace with no per-caller ownership: any caller can see or remove any staged file. Only remove paths you staged in the current task, and treat file management as admin-only on a shared instance — see [Server-Side File Staging](#server-side-file-staging-v1files) below.\n- **No auth by default** — `TALKIES_AUTH_TOKEN` is opt-in; an unconfigured server is wide open to anyone who can reach the port. Require the token and least-privilege network exposure in any non-trusted environment.\n- **Voice cloning requires consent.** Only clone or synthesize a voice you have explicit authorization/consent to use — synthesized speech of a real person can be used for impersonation, fraud, or deception. See [Qwen3-TTS Custom Voices](references/setup.md#qwen3-tts-custom-voices).\n- **Local shell execution.** `references/setup.md` and workflow examples run `docker`, `curl`, and `ffmpeg` directly on the local host — review commands before running them against unfamiliar hosts/images.\n\n## When To Use\n\n- Transcribe audio files (any format ffmpeg decodes — WAV, MP3, M4A, FLAC, OGG, WebM, Opus, MP4 audio).\n- Generate SRT/VTT subtitles for video.\n- Transcribe podcasts, lectures, interviews, voicemails, calls.\n- Stream 16 kHz mono PCM from a microphone or decoded live-audio source through `WS /v1/audio/transcriptions/stream`.\n- Stereo two-mic recordings → per-speaker diarized output (`L:` / `R:` channel tagging).\n- German/French/Spanish ↔ English speech-to-text translation via Canary-1B-Flash.\n- Synthesize speech from text via Kokoro-82M — English (American + British), Spanish, French, Hindi, Italian, Portuguese.\n- Voice-clone speech via Qwen3-TTS-0.6B from a reference `.wav` you provide — drop into `/data/custom-voices/`, immediately appears under `GET /v1/audio/voices` with `origin=custom`. **Only clone voices you're authorized to use, with the speaker's consent** — see [Qwen3-TTS Custom Voices](references/setup.md#qwen3-tts-custom-voices).\n- Voice-clone via `chatterbox-turbo` from the same `/data/custom-voices/` drop — no sibling transcript needed, but the clip must be longer than 5 seconds. Same consent requirement applies.\n- Emotive English delivery via `chatterbox-turbo` — write tags inline in `input`, e.g. `Oh no [sigh] not again.` The tokenizer defines exactly 19: `[angry]` `[fear]` `[surprised]` `[whispering]` `[advertisement]` `[dramatic]` `[narration]` `[crying]` `[happy]` `[sarcastic]` `[clear throat]` `[sigh]` `[shush]` `[cough]` `[groan]` `[sniff]` `[gasp]` `[chuckle]` `[laugh]`. Anything else is read as literal text.\n- Drop-in replacement for `api.openai.com/v1/audio/transcriptions` and `api.openai.com/v1/audio/speech` in existing client code.\n\n## When NOT To Use\n\n- OpenAI-compatible live ASR — `/v1/audio/transcriptions` remains request/response only. Use talkies-specific `WS /v1/audio/transcriptions/stream` for raw PCM. (TTS has one streaming exception: `qwen3-tts-*` + `response_format=pcm` streams chunked PCM — see [Streaming PCM (Qwen3-TTS)](#streaming-pcm-qwen3-tts).)\n- Speaker identification from voice (only stereo-channel diarization is supported, not voice clustering).\n- Per-request `prompt` / `temperature` on `/v1/audio/transcriptions` — accepted for OpenAI compat, **ignored**. (`instructions` on `/v1/audio/speech` is different: Kokoro ignores it, but Qwen3-TTS honors it in most modes — see [Request Body](#request-body).)\n- Japanese / Chinese TTS — Kokoro upstream supports them but talkies filters those voices out (they need the `misaki[ja]` / `misaki[zh]` extras).\n- Kokoro on OpenAI aliases (`alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`) — Kokoro exposes its native voice names only (`af_*`, `bm_*`, etc.). Map client-side. (Qwen3-TTS does ship `alloy` / `echo` / `fable` as builtin voice slugs, but they're voice-cloned samples, not OpenAI's voices — there's no audio compatibility.)\n- `qwen3-tts-0.6b` on CPU — voice cloning hard-fails without CUDA at load time. The `faster_qwen3_tts` upstream raises `ValueError` on non-CUDA devices; talkies surfaces this as a load failure on the first request.\n- `qwen3-tts-0.6b` `speed` parameter — Qwen3-TTS has no playback-rate control. Field is accepted for OpenAI compat but **ignored** (only Kokoro honors `speed`; `chatterbox-turbo` ignores it too).\n- `chatterbox-turbo` for anything but English — it is an English-only checkpoint. Use Kokoro or Qwen3-TTS for other languages.\n- `chatterbox-turbo` on CPU — the slug is registered in the CUDA image only. The model does run on CPU but measures roughly 5-10x slower than realtime, so it is not offered as a CPU slug.\n- `chatterbox-turbo` with default settings where output must be unwatermarked. Every waveform carries Resemble AI's PerTh neural watermark, which the upstream package applies unconditionally. The server can turn it off: set `TALKIES_CHATTERBOX_WATERMARK=false` on the container. That is an operator-side setting, not a per-request one, so a caller cannot change it through the API.\n- `chatterbox-turbo` reference clips of 5 seconds or shorter — rejected with a 400. Supply a longer clip.\n- arm64 hosts — `linux/amd64` only.\n\n## Setup\n\nThe container should already be running. Set the base URL:\n\n```bash\nexport TALKIES_URL=http://localhost:8000\n```\n\nIf the server has `TALKIES_AUTH_TOKEN` set, export it too:\n\n```bash\nexport TALKIES_AUTH_TOKEN=<your-token>\n# every request below needs: -H \"Authorization: Bearer $TALKIES_AUTH_TOKEN\"\n```\n\n**Verify:** `curl $TALKIES_URL/healthz` returns `{\"ok\": true, \"device\": \"...\", \"models\": [...]}`.\n\nFor install / configuration / env vars / CPU vs CUDA images / custom model registry, see [references/setup.md](references/setup.md).\n\n## Quick Start\n\n```bash\n# Discover what's available.\ncurl -s $TALKIES_URL/v1/models | jq\n\n# Simplest transcribe — file upload, JSON response.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq\n\n# Same call, but the audio lives at a URL — talkies downloads + caches it.\n# WARNING: the URL is fetched by the talkies SERVER, not this client, and the\n# download is cached persistently on disk (see \"URL file_path\" below). Don't\n# pass URLs to private/sensitive media unless you trust the server operator.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq\n\n# Full Whisper-shape JSON with per-segment + per-word timestamps.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"response_format=verbose_json\" | jq\n\n# SRT subtitles.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@lecture.mp3\" \\\n  -F \"model=whisper-large-v3\" \\\n  -F \"response_format=srt\" > lecture.srt\n\n# Discover TTS voices, then synthesize an MP3.\ncurl -s $TALKIES_URL/v1/audio/voices | jq\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"kokoro-82m\",\n        \"input\": \"Hello from talkies.\",\n        \"voice\": \"af_heart\",\n        \"response_format\": \"mp3\"\n      }' \\\n  --output hello.mp3\n```\n\n## Live streaming ASR\n\n`WS /v1/audio/transcriptions/stream` is talkies-specific rather than an OpenAI\nendpoint. It accepts headerless 16 kHz, mono, signed 16-bit little-endian PCM.\nSend one JSON `start` message, wait for `ready`, stream binary PCM frames, then\nsend `{\"type\":\"end\"}` for `final` and `stats` (or `{\"type\":\"cancel\"}` to\ndiscard the stream). Use a configured ASR slug with streaming support; native\nsessions are available through parakeet.cpp, Sherpa-ONNX, and Vosk, while\nfaster-whisper uses a bounded rolling window. When authentication is enabled,\nsend the bearer token in the WebSocket upgrade header, never in the URL.\n\nSee the repository's [`docs/streaming.md`](https://github.com/psyb0t/docker-talkies/blob/main/docs/streaming.md)\nfor the exact start/event shapes, a Python client, limits, and custom Sherpa or\nVosk model-registry entries.\n\n## Supported Models\n\n### ASR\n\n| Slug | Family | CPU | CUDA | Languages | Strength |\n|---|---|---|---|---|---|\n| `whisper-large-v3` | faster-whisper | yes | yes | 99 auto-detect | best accuracy, slowest |\n| `whisper-large-v3-turbo` | faster-whisper | yes | yes | 99 auto-detect | sweet spot — fast, accurate |\n| `parakeet-tdt-0.6b-v3` | NeMo TDT | no | yes | English only | very fast on GPU |\n| `nemotron-3.5-asr-0.6b` | parakeet.cpp / ggml | yes | yes | 40+ locales, auto-detect | CPU or CUDA multilingual; pin via `language=` |\n| `canary-180m-flash` | NeMo Canary | yes | yes | English only (small) | smallest, runs anywhere |\n| `canary-1b-flash` | NeMo Canary | no | yes | en/de/fr/es + translation | multilingual, translation |\n| `canary-qwen-2.5b` | NeMo SALM | no | yes | English only | best English accuracy (no timestamps) |\n| `sherpa-zipformer-en-left-64` | Sherpa-ONNX Zipformer | yes | yes | English only | native live ASR, lower left context |\n| `sherpa-zipformer-en-left-128` | Sherpa-ONNX Zipformer | yes | yes | English only | native live ASR, higher left context |\n| `sherpa-zipformer-en-int8-left-64` | Sherpa-ONNX Zipformer INT8 | yes | yes | English only | smaller native live ASR variant |\n| `sherpa-zipformer-en-int8-left-128` | Sherpa-ONNX Zipformer INT8 | yes | yes | English only | smaller native live ASR variant |\n| `vosk-small-en-us-0.15` | Vosk | yes | yes | English only | native live ASR; CPU decoder |\n| `wav2vec2-xlsr-53-espeak` | wav2vec2 CTC phoneme | yes | yes | multilingual | IPA phones, not words; eSpeak alphabet |\n| `zipa-ipa` | ZIPA (Zipformer CTC phoneme) | yes | yes | multilingual | IPA phones, not words; 71 MB, very fast |\n\nPick by use case:\n- **General-purpose:** `whisper-large-v3-turbo`.\n- **English-only, max accuracy on GPU:** `canary-qwen-2.5b` (but no per-segment timestamps).\n- **Translation EN↔DE/FR/ES:** `canary-1b-flash` (requires custom model registry — see [Translation](#translation)).\n- **Low-overhead live English ASR:** one of the Sherpa Zipformer variants or `vosk-small-en-us-0.15`; Sherpa uses CUDA in the CUDA image, while Vosk remains CPU-decoded.\n- **Phones, not words (pronunciation checks, phonetics):** `wav2vec2-xlsr-53-espeak` or `zipa-ipa`. Both return a space-separated IPA stream with per-phone timestamps and run no language model, so mispronunciations are not corrected away. `zipa-ipa` is the small fast pick; `wav2vec2-xlsr-53-espeak` is the larger transformers-native one.\n\n### TTS\n\n3 engines / 4 backends across 8 slugs. Kokoro ships in two runtimes (`kokoro-82m` PyTorch, `kokoro-82m-nvidia` ONNX/ORT) — same weights, same voice catalog, same wire format. Qwen3-TTS ships 5 CUDA-only slugs across three modes (base cloning / custom_voice preset speakers / voice_design). Mode is implicit in the slug — see [Qwen3-TTS Modes](#qwen3-tts-modes). `chatterbox-turbo` (CUDA-only, English) rounds out the set — 19 inline emotion tags, transcript-free voice cloning; see [When To Use](#when-to-use) above.\n\n| Slug | Family | Mode | CPU | CUDA | Languages | Voices |\n|---|---|---|---|---|---|---|\n| `kokoro-82m` | Kokoro (PyTorch in-process, 24 kHz) | — | yes | yes | en (US + UK), es, fr, hi, it, pt | 41 baked (discover via `GET /v1/audio/voices`) |\n| `kokoro-82m-nvidia` | Kokoro (ONNX via ORT, 24 kHz) | — | yes | yes | en (US + UK), es, fr, hi, it, pt | 41 baked (same catalog as `kokoro-82m`) |\n| `qwen3-tts-0.6b` | Qwen3-TTS (24 kHz) | base | no | yes | 17 (en, zh, ja, ko, fr, de, es, it, pt, ru, vi, th, id, ar, tr, pl, nl) | 3 builtin samples + any `.wav` under `/data/custom-voices/` |\n| `qwen3-tts-1.7b` | Qwen3-TTS (24 kHz) | base | no | yes | 10 (en, zh, ja, ko, fr, de, es, it, pt, ru) | 3 builtin samples + any `.wav` under `/data/custom-voices/` |\n| `qwen3-tts-0.6b-custom` | Qwen3-TTS (24 kHz) | custom_voice | no | yes | en, zh, ja, ko | 9 preset speakers (`instructions` dropped — 0.6B limitation) |\n| `qwen3-tts-1.7b-custom` | Qwen3-TTS (24 kHz) | custom_voice | no | yes | en, zh, ja, ko | 9 preset speakers + emotion via `instructions` |\n| `qwen3-tts-1.7b-design` | Qwen3-TTS (24 kHz) | voice_design | no | yes | en, zh, ja, ko | voice synthesized from NL description in `instructions` (required) |\n| `chatterbox-turbo` | Chatterbox Turbo (24 kHz, buffered only) | — | no | yes | English only | `builtin` speaker, or any `.wav` (>5s) under `/data/custom-voices/` |\n\nPick by use case:\n- **General-purpose multi-voice TTS:** `kokoro-82m` — fast, 41 baked voices, runs on CPU. Use `kokoro-82m-nvidia` for the ONNX/ORT execution path (CUDA EP on the CUDA image, CPU EP otherwise).\n- **Voice cloning from a reference clip:** `qwen3-tts-0.6b` / `qwen3-tts-1.7b` — drop a `.wav` into `/data/custom-voices/`, immediately usable. CUDA required.\n- **Preset speakers (no reference WAV):** `qwen3-tts-0.6b-custom` / `qwen3-tts-1.7b-custom` — 9 baked speakers; the 1.7B honours `instructions` for emotion. CUDA required.\n- **Invent a voice from a description:** `qwen3-tts-1.7b-design` — the NL description goes in `instructions`. CUDA required.\n- **Expressive English delivery, transcript-free cloning:** `chatterbox-turbo` — 19 inline emotion tags in `input`, clones from a bare `.wav` (no sibling transcript needed, clip must be longer than 5 seconds). English only, CUDA required, output always carries a neural watermark.\n\n`canary-qwen-2.5b` produces no segment/word timestamps — `verbose_json.segments` and `.words` come back empty, `srt`/`vtt` collapse to a single full-duration cue. Transcription itself is whole-file. Use a Whisper or Canary multitask slug if you need timing.\n\n## API — `POST /v1/audio/transcriptions`\n\nMultipart form. Same field names as OpenAI's transcription endpoint where they overlap.\n\n### Request Fields\n\n| Field | Required | Default | Notes |\n|---|---|---|---|\n| `file` | one of `file`/`file_path` | — | Audio file. Capped at `TALKIES_MAX_UPLOAD_BYTES` (default 100 MB). |\n| `file_path` | one of `file`/`file_path` | — | Either a path under the staging area (`/v1/files`) or an `http(s)://` URL (downloaded + cached server-side). Not subject to the 100 MB upload cap; URL downloads capped by `TALKIES_MAX_DOWNLOAD_BYTES` (default 1 GiB). |\n| `model` | yes | — | One of the configured slugs (see `GET /v1/models`). Unknown → 404. |\n| `language` | no | model default | ISO-639-1 code. Whisper auto-detects when omitted; Canary uses its `default_source_lang`. |\n| `response_format` | no | `json` | `json` / `text` / `verbose_json` / `srt` / `vtt`. |\n| `timestamp_granularities[]` | no | — | Accepted for OpenAI compat; ignored — `verbose_json` always emits both segment + word. |\n| `prompt` | no | — | **Accepted, ignored.** |\n| `temperature` | no | — | **Accepted, ignored.** |\n| `diarization` | no | `false` | Stereo-channel diarization. Requires 2-channel input — mono returns 400. |\n\nExactly one of `file` or `file_path` must be set — passing both or neither returns 400.\n\n### Response Formats\n\n| `response_format` | Content-Type | Shape |\n|---|---|---|\n| `json` (default) | `application/json` | `{\"text\": \"...\"}` — just the transcript. |\n| `text` | `text/plain` | The transcript as plain text. |\n| `verbose_json` | `application/json` | Full Whisper shape — `task`, `language`, `duration`, `text`, `segments[]`, `words[]`. |\n| `srt` | `application/x-subrip` | SubRip subtitle file, one cue per VAD-segmented chunk. |\n| `vtt` | `text/vtt` | WebVTT subtitle file, one cue per VAD-segmented chunk. |\n\n`json` shape:\n```json\n{ \"text\": \" full transcript as a single string\" }\n```\n\n`verbose_json` shape — `segments` and `words` are always present (empty arrays for backends with no alignment output):\n```json\n{\n  \"task\": \"transcribe\",\n  \"language\": \"en\",\n  \"duration\": 6.42,\n  \"text\": \" full transcript\",\n  \"segments\": [{ \"id\": 0, \"start\": 0.0, \"end\": 2.31, \"text\": \" ...\", \"tokens\": [], \"temperature\": 0.0, \"avg_logprob\": null, \"compression_ratio\": null, \"no_speech_prob\": null }],\n  \"words\": [{ \"word\": \" the\", \"start\": 0.0, \"end\": 0.12 }]\n}\n```\n\nWhisper-only confidence fields (`avg_logprob`, `compression_ratio`, `no_speech_prob`) are emitted as `null` regardless of backend so clients reading them don't crash. `tokens` is always `[]`.\n\nThe `sherpa` and `vosk` executors add a per-word `confidence` in the 0–1 range to each entry in `words` — Vosk reports the decoder's own score, Sherpa derives one from the model's per-token acoustic log-probabilities. No other backend emits it, so treat the field as optional.\n\n### Stereo Diarization\n\nPass `diarization=true` and upload a 2-channel file. Left channel = speaker `L`, right channel = speaker `R`. Each channel is transcribed independently, the two timelines are merged chronologically by segment start time.\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@interview-stereo.wav\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"diarization=true\" \\\n  -F \"response_format=verbose_json\" | jq\n```\n\nWhat changes:\n- `verbose_json` — every segment/word gets `\"channel\": \"L\"` or `\"R\"`. Segments re-numbered after merge.\n- `text` / `response_format=text` — rebuilt as alternating turn lines: `L: ...\\nR: ...\\n...`. Consecutive same-channel segments collapsed into one line per turn.\n- `srt` / `vtt` — each cue prefixed with `L:` / `R:`.\n\nCaveats:\n- Exactly **2 channels** required. Mono → 400. >2 channels → 400.\n- Latency ~2× the mono case (model runs sequentially on each channel).\n- The technique is exact for true two-mic setups (interview rigs, podcast splits). It does NOT magically separate speakers from a single-mic recording that's been rendered to stereo.\n\n### Translation\n\nCanary multitask models can translate speech → text in a non-source language. `canary-1b-flash` covers en↔de, en↔fr, en↔es. **The task is baked into the model slug**, not passed per-request — you add a translation-specific slug via custom `models.json` (see [Customizing the model registry](references/setup.md#customizing-the-model-registry)):\n\n```json\n{\n  \"models\": {\n    \"canary-1b-flash-de2en\": {\n      \"repo\": \"nvidia/canary-1b-flash\",\n      \"executor\": \"canary_multitask\",\n      \"default_source_lang\": \"de\",\n      \"default_target_lang\": \"en\",\n      \"default_task\": \"s2t_translation\",\n      \"languages\": [\"de\"]\n    }\n  }\n}\n```\n\nThen call it normally — `text` carries the English translation:\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@german-clip.wav\" \\\n  -F \"model=canary-1b-flash-de2en\" | jq\n```\n\n`canary-180m-flash` is English-ASR-only — don't point a translation slug at it. `canary-qwen-2.5b` is English ASR only too.\n\n### Long Files + VAD Chunking\n\nAudio longer than 30 s (`TALKIES_VAD_CHUNK_THRESHOLD`) gets sliced through Silero VAD into ≤28 s speech regions before being handed to the backend. Timestamps are re-assembled by offsetting each chunk's segment/word timings — you get one continuous `segments` list spanning the whole file.\n\nNo client-side change. Long files just work. Verify by checking `duration` in `verbose_json`.\n\n### Error Contract\n\n| Status | Shape | When |\n|---|---|---|\n| 200 | per `response_format` | success |\n| 400 | `{\"detail\": \"...\"}` | bad audio, mono+diarization, >2 ch+diarization, both/neither of `file`/`file_path`, invalid file_path, URL download failure (DNS, HTTP error, size exceeded, SSRF blocked) |\n| 401 | `{\"detail\": \"...\"}` | only when `TALKIES_AUTH_TOKEN` is set: missing/wrong bearer. Includes `WWW-Authenticate: Bearer`. |\n| 404 | `{\"detail\": \"...\"}` | unknown model slug, `file_path` references missing file, model-evict on an unloaded model, file op on a missing `/v1/files` path |\n| 413 | `{\"detail\": \"...\"}` | upload exceeded `TALKIES_MAX_UPLOAD_BYTES` (multipart `file` and `PUT /v1/files/{path}` only — not `file_path` URL) |\n| 422 | `{\"detail\": [...]}` | Pydantic validation (missing fields, wrong types) |\n| 500 | `{\"detail\": \"...\"}` | unhandled backend failure |\n\n## API — `POST /v1/audio/speech` (TTS)\n\nJSON body (not multipart). Returns the encoded audio bytes in the body with the matching `Content-Type` — no JSON envelope.\n\n**Every call sends your `input` text (and, for voice-cloning slugs, the referenced voice sample) to whatever `$TALKIES_URL` points at — that data leaves your host.** Point `$TALKIES_URL` only at a talkies instance you run or explicitly trust; prefer HTTPS when it's not localhost/LAN. Don't synthesize sensitive or confidential text through a server you don't control.\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"kokoro-82m\",\n        \"input\": \"The quick brown fox jumps over the lazy dog.\",\n        \"voice\": \"af_heart\",\n        \"response_format\": \"mp3\",\n        \"speed\": 1.0\n      }' \\\n  --output fox.mp3\n```\n\n### Request Body\n\n| Field | Required | Default | Notes |\n|---|---|---|---|\n| `model` | yes | — | TTS model slug. Kokoro: `kokoro-82m`, `kokoro-82m-nvidia`. Qwen3-TTS: `qwen3-tts-0.6b`, `qwen3-tts-1.7b` (base/cloning), `qwen3-tts-0.6b-custom`, `qwen3-tts-1.7b-custom` (preset speakers), `qwen3-tts-1.7b-design` (voice from NL description). Chatterbox: `chatterbox-turbo` (English only). Unknown → 404. ASR slug → 400. |\n| `input` | yes | — | Text to synthesize. Empty / whitespace-only → 400. No fixed length cap; for very long inputs split client-side. |\n| `voice` | no | model `default_voice` | Semantics shift per Qwen3 mode (see [Qwen3-TTS Modes](#qwen3-tts-modes)). Kokoro: voice name (default `af_heart`). Qwen3 `base`: path of a reference WAV (default `alloy`). Qwen3 `custom_voice`: one of the 9 preset speakers (default `Vivian`). Qwen3 `voice_design`: ignored — sentinel `\"design\"`. `chatterbox-turbo`: `builtin` (default) or the name of a `.wav` under `/data/custom-voices/`, longer than 5 seconds — shorter clips → 400. Unknown → 400 with catalog listed. |\n| `response_format` | no | `mp3` | `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm`. |\n| `speed` | no | `1.0` | Playback rate, Kokoro only. Clamped to `[0.25, 4.0]`. **Ignored** by every Qwen3-TTS slug (no speed control in Qwen3-TTS) and by `chatterbox-turbo`. |\n| `instructions` | no | — | Free-form style prompt. **Required** for `qwen3-tts-1.7b-design` (the NL voice description; empty → 400). **Honoured** by Qwen3-TTS `base` mode and `qwen3-tts-1.7b-custom` (threaded as `instruct`). **Dropped** by `qwen3-tts-0.6b-custom` (0.6B CustomVoice checkpoint limitation — logs a WARNING) and both Kokoro slugs (no instruction input). Accepted on every slug for OpenAI parity. |\n| `language` | no | model `default_language` (`English`) | **Non-OpenAI extra field** (send via `extra_body={\"language\": \"...\"}` on official SDKs). Selects the spoken language for Qwen3 `custom_voice` / `voice_design`; `base` mode reads it from the voice's sibling `.lang` file. Silently ignored by Kokoro. |\n| `temperature` | no | `0.9` | **Non-OpenAI extra, Qwen3-TTS only** (`extra_body`). Sampler temperature, `[0.0, 2.0]`. Ignored by Kokoro. |\n| `top_k` | no | `50` | **Non-OpenAI extra, Qwen3-TTS only.** Top-k truncation, `[1, 1000]`. Ignored by Kokoro. |\n| `top_p` | no | `1.0` | **Non-OpenAI extra, Qwen3-TTS only.** Nucleus sampling, `[0.0, 1.0]`. Ignored by Kokoro. |\n| `repetition_penalty` | no | `1.05` | **Non-OpenAI extra, Qwen3-TTS only.** Penalizes codec-token repeats, `[0.5, 2.0]`. Ignored by Kokoro. |\n| `max_new_tokens` | no | `2048` (model max) | **Non-OpenAI extra, Qwen3-TTS only.** Codec-step cap, `[1, 2048]`. Ignored by Kokoro. |\n| `do_sample` | no | `true` | **Non-OpenAI extra, Qwen3-TTS only.** `false` = greedy decode. Ignored by Kokoro. |\n\nOut-of-range sampling values → 422 (Pydantic validation).\n\n### Qwen3-TTS Modes\n\nThe Qwen3-TTS mode is implicit in the model slug — the OpenAI wire format stays pure (`model` / `voice` / `instructions` / `input`), with `voice` and `instructions` carrying mode-specific semantics. No new endpoints.\n\n| Mode | Slugs | What `voice` means | What `instructions` means |\n|---|---|---|---|\n| **base** (voice cloning) | `qwen3-tts-0.6b`, `qwen3-tts-1.7b` | Path of a reference `.wav` under the voices dirs (`.wav` stripped) | Optional style hint (passed as `instruct`) |\n| **custom_voice** (preset speakers) | `qwen3-tts-0.6b-custom`, `qwen3-tts-1.7b-custom` | One of 9 preset speaker names | Emotion / style cue — 1.7B honours it; 0.6B drops it (checkpoint limitation, logs a WARNING) |\n| **voice_design** (NL description) | `qwen3-tts-1.7b-design` | Ignored — sentinel `\"design\"` | **Required.** NL description of the voice (e.g. \"A warm, friendly young female voice\"). Empty → 400. |\n\nThe 9 `custom_voice` preset speakers (also returned by `GET /v1/audio/voices` for the chosen slug): `Vivian`, `Serena`, `Uncle_Fu`, `Dylan`, `Eric` (Chinese), `Ryan`, `Aiden` (English), `Ono_Anna` (Japanese), `Sohee` (Korean).\n\n### Output Formats\n\n`response_format` picks the encoder applied to Kokoro's raw 24 kHz mono PCM. ffmpeg does the conversion in-process; no temp files.\n\n| `response_format` | Content-Type | Codec / container | Notes |\n|---|---|---|---|\n| `mp3` (default) | `audio/mpeg` | libmp3lame, 128 kbps CBR | Most universal. |\n| `opus` | `audio/ogg` | libopus, 64 kbps VBR, Ogg container | Best quality-per-byte for speech. |\n| `aac` | `audio/aac` | AAC-LC, 128 kbps, ADTS | iOS-friendly. |\n| `flac` | `audio/flac` | FLAC | Lossless. |\n| `wav` | `audio/wav` | PCM s16le, 24 kHz mono, RIFF header | Lossless, largest. |\n| `pcm` | `application/octet-stream` | Raw PCM s16le, 24 kHz mono — no container, no header | Real-time chaining. Caller must know sample rate / format. |\n\n### Streaming PCM (Qwen3-TTS)\n\n`qwen3-tts-*` slugs stream `response_format=pcm` requests as chunked int16 PCM instead of buffering the whole utterance — first-audio latency drops from seconds to sub-second. Kokoro and every non-`pcm` format are unaffected (buffered as normal). Response carries an `X-Sample-Rate` header; chunk size (codec steps per yielded chunk) is tunable via `TALKIES_QWEN3_STREAM_CHUNK_SIZE` (default 8) — see [references/setup.md](references/setup.md).\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"qwen3-tts-0.6b\",\"input\":\"Streaming test.\",\"response_format\":\"pcm\"}' \\\n  --output stream.pcm\n```\n\n### Voices\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/voices | jq\n```\n\nReturns `{\"voices\": [{\"voice\", \"model\", \"default\", \"origin\"}]}`. The `origin` field is only present for engines that distinguish baked-in vs user-supplied voices (currently `qwen3-tts-0.6b` — `\"builtin\"` for image-baked samples, `\"custom\"` for `/data/custom-voices/` mounts). Kokoro entries omit `origin`.\n\n**Kokoro voices** encode `<lang_code><gender>_<name>`:\n\n| Prefix | Language |\n|---|---|\n| `af_` / `am_` | American English (female / male) |\n| `bf_` / `bm_` | British English (female / male) |\n| `ef_` / `em_` | Spanish |\n| `ff_` | French |\n| `hf_` / `hm_` | Hindi |\n| `if_` / `im_` | Italian |\n| `pf_` / `pm_` | Portuguese (Brazilian) |\n\n41 voices ship in the image. Japanese (`jf_*` / `jm_*`) and Chinese (`zf_*` / `zm_*`) are filtered out because they need the optional `misaki[ja]` / `misaki[zh]` extras (MeCab + pypinyin chains).\n\n**Qwen3-TTS voices** come from two on-disk dirs merged into one catalog:\n\n- `/opt/talkies/qwen3-voices/` — baked into the CUDA image. Ships three curated samples (`alloy`, `echo`, `fable`) so voice cloning works out-of-the-box. `origin=builtin`.\n- `/data/custom-voices/` — host-mounted via the data volume. Drop `foo/bar/me.wav` and voice `foo/bar/me` immediately appears in `GET /v1/audio/voices` (catalog is rescanned per request — no restart). `origin=custom`.\n\nVoice names are the wav's path relative to its parent dir with `.wav` stripped — nested subdirs are preserved. `custom-voices/team-a/jane.wav` → voice `team-a/jane`. Custom voices shadow builtin voices with the same name; dropping a `custom-voices/alloy.wav` overrides the builtin `alloy` sample (its `origin` flips to `custom`).\n\nOptional sibling metadata next to each `<name>.wav`:\n- `<name>.txt` — reference transcript for the clip (ICL voice cloning works without it, but clone fidelity is noticeably better with a faithful transcript).\n- `<name>.lang` — language label string (defaults to `English`).\n\nPath-traversal guard: hostile symlinks whose `resolve()` escapes the voices dir are skipped (the wav can't be used to read arbitrary host files as a voice prompt).\n\n```bash\n# Add a custom clone voice (server picks it up on next request — no restart).\nmkdir -p ~/talkies-data/custom-voices/team-a\ncp jane-reading.wav ~/talkies-data/custom-voices/team-a/jane.wav\necho \"And the silken sad uncertain rustling of each purple curtain.\" \\\n  > ~/talkies-data/custom-voices/team-a/jane.txt\n\n# Use it.\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"qwen3-tts-0.6b\",\n        \"input\": \"Hello from a cloned voice.\",\n        \"voice\": \"team-a/jane\",\n        \"response_format\": \"wav\"\n      }' \\\n  --output cloned.wav\n```\n\n**First synth is slow** on Qwen3-TTS — the predictor + talker CUDA graphs are captured on first call (~30-60 s on a mid-range GPU). Subsequent generations are sub-second. The model and graphs stay resident until evicted by sibling load or the idle sweeper.\n\n### Error Contract (TTS)\n\n| Status | When |\n|---|---|\n| 200 | success (audio bytes in body) |\n| 400 | empty `input`, unknown `voice`, unsupported `response_format`, model isn't TTS (e.g. POSTing `whisper-large-v3` here) |\n| 401 | `TALKIES_AUTH_TOKEN` set, missing / wrong bearer |\n| 404 | unknown `model` slug |\n| 422 | Pydantic validation (missing required fields, wrong types) |\n| 500 | unhandled ffmpeg or kokoro internal failure |\n| 503 | TTS snapshot files missing under `${TALKIES_DATA_DIR}/models/<slug>/` (slug excluded from `TALKIES_ENABLED_MODELS` but still being called); or `qwen3-tts-0.6b` requested on a non-CUDA device (the backend hard-fails at load time) |\n\n## Resource-Management Endpoints (Ollama-Style)\n\ntalkies mirrors a subset of [speaches](https://github.com/speaches-ai/speaches) / Ollama, so a LiteLLM proxy can drive both.\n\n| Endpoint | Behavior |\n|---|---|\n| `GET /healthz` | Unauthenticated liveness. Returns `{ok, device, models}`. |\n| `GET /v1/models` | OpenAI-style list of configured slugs. Each entry includes a `modality` field (`asr` or `tts`) so clients can filter. |\n| `GET /api/ps` | Currently-loaded models with per-model `idle_seconds`. |\n| `DELETE /api/ps/{model_id}` | Evict one model from memory. Slug can be URL-encoded (`/` → `%2F`). 404 if not loaded. |\n| `POST /unload` | Evict every loaded model. Returns the list actually unloaded. |\n\nModel eviction (`DELETE /api/ps/...`, `POST /unload`) forces a cold-load for anyone mid-request, so only do it for explicit maintenance the user asked for (e.g. \"free up VRAM\"). Auth is only enforced when `TALKIES_AUTH_TOKEN` is set — require it on shared deployments and don't expose these routes on an unauthenticated network.\n\nBehind these: an **idle sweeper** runs every `TALKIES_SWEEPER_INTERVAL` s (default 60) and unloads anything not used in `TALKIES_MODEL_TTL` s (default 600). Set `TALKIES_MODEL_TTL=0` to disable.\n\nThere's also **sibling eviction at request time** — every transcribe or speech request evicts other loaded models so VRAM doesn't get split. ASR and TTS share the same pool; loading Kokoro evicts a resident Whisper and vice versa. One model resident at a time, per container. If you need two models simultaneously, run two containers.\n\n```bash\n# Which models are loaded right now.\ncurl -s $TALKIES_URL/api/ps | jq\n\n# Free VRAM after a job — evict one model.\ncurl -s -X DELETE \"$TALKIES_URL/api/ps/whisper-large-v3-turbo\"\n\n# Or evict everything.\ncurl -s -X POST $TALKIES_URL/unload | jq\n```\n\n## Server-Side File Staging (`/v1/files`)\n\nFor repeated transcribes of the same file (different `response_format`, different model, iterating on params), stage the file once and reference it by path. Files land under `${TALKIES_DATA_DIR}/files/<path>`.\n\n**Staged files persist until explicitly removed** — nothing auto-expires them — and `GET /v1/files` enumerates every staged path to anyone who can reach the API. Don't stage sensitive/private media on a server without auth enabled (`TALKIES_AUTH_TOKEN`); clean up staged files when done.\n\n**Guardrail — this is a shared, unisolated bucket, not a private workspace.** There's no per-caller ownership: any path any caller staged is listable and readable by any other caller with API access. An agent must:\n- only read or delete paths it staged itself in the current, user-approved workflow;\n- never call `GET /v1/files` to browse/enumerate what other callers have staged, and never delete a path it didn't create;\n- clean up what it staged once the workflow is done, since nothing expires automatically.\n\nAt the deployment level: require `TALKIES_AUTH_TOKEN` by default, treat per-caller isolation and retention limits as the operator's responsibility (talkies itself provides neither).\n\n| Endpoint | Behavior |\n|---|---|\n| `GET /v1/files` | List every staged file. Returns `{\"files\": [{\"path\", \"size\", \"modified\"}]}`. **Enumerable by anyone with API access — no per-file ownership/isolation.** |\n| `PUT /v1/files/{path}` | Upload raw bytes (`--data-binary @local-file`). Capped at `TALKIES_MAX_UPLOAD_BYTES`. Atomic write (`.part` → rename). |\n| `GET /v1/files/{path}` | Streams file back. Content-Type guessed by extension. 404 if missing. |\n| `DELETE /v1/files/{path}` | Removes a staged file and prunes empty parent dirs (404 if missing). Files don't self-expire, so call this to clean up when done. On a shared bucket every path is listable by any caller, so only remove paths you staged yourself. |\n\n```bash\n# Stage once.\ncurl -X PUT --data-binary @lecture.mp3 \\\n  -H \"Content-Type: audio/mpeg\" \\\n  $TALKIES_URL/v1/files/lectures/2026-03-15/lecture.mp3\n\n# Reuse across multiple transcribe calls.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=lectures/2026-03-15/lecture.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"response_format=verbose_json\" | jq\n\n# Cleanup.\ncurl -X DELETE $TALKIES_URL/v1/files/lectures/2026-03-15/lecture.mp3\n```\n\nPath safety: null bytes, backslashes, `.` / `..` segments and double slashes are rejected (400). Symlinks pointing outside the root are refused. Leading `/` is stripped — `/foo/bar.mp3` and `foo/bar.mp3` resolve identically.\n\n### URL `file_path` (Download + Cache)\n\n`file_path` also accepts `http://` / `https://` URLs. First request downloads to `${TALKIES_DATA_DIR}/files/downloads/<sha256(url)[:16]>-<basename>`, subsequent requests with the same URL hit the cache.\n\n**The download happens server-side, and the result is cached persistently on the talkies server's disk** — not a transient client-side fetch. Anyone who can reach the API can later list/read that cached copy via `GET /v1/files` (see [Server-Side File Staging](#server-side-file-staging-v1files) above). Don't pass URLs to private/sensitive media unless the talkies server itself is trusted and access-controlled (`TALKIES_AUTH_TOKEN`); invalidate the cache entry when done.\n\n```bash\n# First call: downloads, transcribes off the cached copy.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq\n\n# Second call: same URL → cache hit, no re-download.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=canary-1b-flash\" \\\n  -F \"response_format=srt\" > ep-042.srt\n```\n\nDownloads appear in `GET /v1/files` listings under `downloads/`. Invalidate a single cached URL by removing it from `/v1/files/downloads/`.\n\nConstraints applied during download:\n- Size capped by `TALKIES_MAX_DOWNLOAD_BYTES` (default 1 GiB).\n- 5 redirect hops max; SSRF guard re-applied at every hop.\n- 10 s connect, 300 s per-chunk read timeout.\n- SSRF off by default. Set `TALKIES_BLOCK_PRIVATE_DOWNLOADS=true` to reject URLs whose hostname resolves to private/loopback/link-local/multicast/reserved IPs.\n\n## MCP Endpoint (`/v1/mcp`)\n\ntalkies exposes a [Model Context Protocol](https://modelcontextprotocol.io) server over Streamable HTTP at `/v1/mcp`. Same FastAPI process, same `BACKENDS` / `REGISTRY`, same auth middleware — a model loaded by the MCP `transcribe` tool is the same instance the HTTP endpoint sees.\n\nMCP exposes the ASR surface only. TTS (`/v1/audio/speech`) is HTTP-only — generated audio bytes don't round-trip through JSON-RPC cleanly. `list_models` filters out TTS slugs so `transcribe` only ever sees ASR backends.\n\n| Tool | What it does |\n|---|---|\n| `list_models` | Discover ASR slugs (TTS slugs are filtered out). Returns `[{slug, executor, default_source_lang, default_target_lang, default_task, loaded}]`. |\n| `transcribe` | Run ASR on a `file_path` (URL or staged path). Args: `model`, `language?`, `response_format?` (`json`/`verbose_json`/`text`/`srt`/`vtt`), `diarization?`. JSON formats return a JSON-encoded string; text/srt/vtt return raw. |\n| `list_files` | Same payload as `GET /v1/files`. |\n| `put_file` | Upload to staging. Body is base64 (`content_base64`). Decoded size capped at `TALKIES_MAX_UPLOAD_BYTES`. **For big files, prefer `PUT /v1/files/{path}` over HTTP** — JSON-RPC + base64 chews token budget. |\n| `get_file` | Read a staged file as base64. Same size cap. Same advice — for big bytes, hit `GET /v1/files/{path}` over HTTP. |\n| `delete_file` | Remove a staged file, prune empty parents. |\n\nThe transport requires `Accept: application/json, text/event-stream`. Wire it into Claude Code:\n\n```bash\nclaude mcp add --transport http talkies $TALKIES_URL/v1/mcp\n```\n\nWith auth:\n\n```bash\nclaude mcp add --transport http talkies $TALKIES_URL/v1/mcp \\\n  --header \"Authorization: Bearer $TALKIES_AUTH_TOKEN\"\n```\n\nNote: the canonical mount path is `/v1/mcp/` (trailing slash). Bare `/v1/mcp` is rewritten internally to `/v1/mcp/` so clients that don't follow Starlette's 307 redirect work too.\n\n### Raw JSON-RPC\n\nFor debugging or non-MCP-aware callers, hit it as JSON-RPC over HTTP POST:\n\n```bash\n# tools/list\ncurl -s $TALKIES_URL/v1/mcp/ \\\n  -H \"Content-Type: application/json\" \\\n  -H \"Accept: application/json, text/event-stream\" \\\n  -d '{\"jsonrpc\": \"2.0\", \"id\": 1, \"method\": \"tools/list\"}'\n\n# tools/call\ncurl -s $TALKIES_URL/v1/mcp/ \\\n  -H \"Content-Type: application/json\" \\\n  -H \"Accept: application/json, text/event-stream\" \\\n  -d '{\n    \"jsonrpc\": \"2.0\", \"id\": 2, \"method\": \"tools/call\",\n    \"params\": {\n      \"name\": \"transcribe\",\n      \"arguments\": {\n        \"file_path\": \"https://example.com/clip.mp3\",\n        \"model\": \"whisper-large-v3-turbo\",\n        \"response_format\": \"json\"\n      }\n    }\n  }'\n```\n\n## Bearer-Token Auth\n\nIf `TALKIES_AUTH_TOKEN` is set on the server, every route except `/healthz` and CORS preflight (`OPTIONS`) requires `Authorization: Bearer <token>`. Wrong/missing token returns 401 with `WWW-Authenticate: Bearer`. Compared with `hmac.compare_digest` (constant-time).\n\n```bash\ncurl -H \"Authorization: Bearer $TALKIES_AUTH_TOKEN\" $TALKIES_URL/v1/models\n```\n\nEmpty / unset token = wide open. For untrusted networks, combine the token with a reverse proxy doing TLS + rate limiting.\n\n## Typical Workflows\n\n### Quick one-off transcribe\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq -r .text\n```\n\n### Generate subtitles for a video\n\n```bash\nffmpeg -i video.mp4 -vn -acodec libmp3lame audio.mp3\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3\" \\\n  -F \"response_format=srt\" > video.srt\n# burn in:  ffmpeg -i video.mp4 -vf subtitles=video.srt -c:a copy video-subbed.mp4\n```\n\n### Iterate on the same file with different settings\n\n```bash\n# Stage once.\ncurl -X PUT --data-binary @lecture.mp3 \\\n  -H \"Content-Type: audio/mpeg\" \\\n  $TALKIES_URL/v1/files/work/lecture.mp3\n\n# Try different models / formats without re-uploading.\nfor fmt in json verbose_json srt; do\n  curl -s $TALKIES_URL/v1/audio/transcriptions \\\n    -F \"file_path=work/lecture.mp3\" \\\n    -F \"model=whisper-large-v3-turbo\" \\\n    -F \"response_format=$fmt\" > \"lecture.$fmt\"\ndone\n\n# Clean up the staged file when done (see the /v1/files reference above).\n```\n\n### Diarized interview transcript\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@interview-stereo.wav\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"diarization=true\" \\\n  -F \"response_format=text\"\n# stdout:\n#   L: hi how's it going\n#   R: not bad you\n#   L: cool man\n```\n\n### Synthesize speech from text\n\n```bash\n# Default voice, MP3 output.\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"kokoro-82m\",\"input\":\"Greetings, human.\"}' \\\n  --output greetings.mp3\n\n# Pick a voice from GET /v1/audio/voices, choose a format.\ncurl -s $TALKIES_URL/v1/audio/voices | jq -r '.voices[].voice'\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"kokoro-82m\",\n        \"input\": \"Buongiorno, mondo.\",\n        \"voice\": \"if_sara\",\n        \"response_format\": \"opus\"\n      }' \\\n  --output ciao.opus\n```\n\n### Free VRAM after a job\n\n```bash\ncurl -s -X POST $TALKIES_URL/unload | jq\n```\n\n### Bulk transcribe from URLs\n\n```bash\nfor url in $(cat urls.txt); do\n  curl -s $TALKIES_URL/v1/audio/transcriptions \\\n    -F \"file_path=$url\" \\\n    -F \"model=whisper-large-v3-turbo\" \\\n    -F \"response_format=text\"\n  echo \"---\"\ndone\n```\n\nThe first hit on each URL downloads + caches; re-running the loop is free.\nFor local files, replace the `file_path` form field with `file=@path/to/audio`\nin the same request shape.\n\n## Tips\n\n1. **Use `whisper-large-v3-turbo`** as your default — it's the speed/quality sweet spot for general-purpose ASR. Switch to `whisper-large-v3` only when you need the last few % of accuracy on hard audio.\n2. **URL `file_path` over multipart upload** — if the audio is already at a URL, send the URL. Saves bandwidth (the file isn't going up and then back down), gets cached server-side, no upload size cap.\n3. **Stage repeated files** via `PUT /v1/files/{path}` and call with `file_path=` to avoid re-uploading on every retry/iteration.\n4. **`response_format=text`** for the \"just give me the string\" case — no `jq -r .text` needed, content-type is `text/plain`.\n5. **One active model family at a time** — every transcribe request evicts other loaded models. Multiple live streams may share one pinned model up to `TALKIES_STREAM_MAX_CONNECTIONS`; a request or stream for a different model receives a conflict until the streams end. Use two containers if you need concurrent models.\n6. **`POST /unload` after a job** — explicit eviction frees VRAM/RAM faster than waiting for the 10-min idle sweeper. Useful in CI / batch scripts.\n7. **`canary-qwen-2.5b` has no timestamps** — `verbose_json.segments` / `.words` come back empty, `srt`/`vtt` collapse to one cue. Use a Whisper or Canary multitask slug if you need timing data.\n8. **Diarization requires true stereo** — if your \"stereo\" file is the same mono signal copied to both channels, diarization won't separate speakers. The technique is exact for two-mic setups, useless otherwise.\n9. **Long files just work** — VAD chunking happens transparently. Don't pre-split. Send the whole file.\n10. **ASR's `prompt` / `temperature` are ignored** even though the request schema accepts them. TTS's `instructions` is different — Kokoro ignores it, but Qwen3-TTS honors it in `base` mode and `*-custom` (except the 0.6B checkpoint) and requires it for `*-design`.\n11. **Watch `/api/ps`** to see what's resident. A request that hangs at \"loading model\" is doing the first cold load — subsequent calls are fast.\n12. **Customizing the model registry** for translation slugs or to restrict the served set — see [references/setup.md](references/setup.md#customizing-the-model-registry).\n13. **Kokoro uses native voice names** — no OpenAI aliases. Hit `GET /v1/audio/voices` once to discover what's shipped; pass the `voice` field accordingly. The 41 voices cover en (US + UK), es, fr, hi, it, pt; ja/zh are filtered out.\n14. **Voice cloning is `qwen3-tts-0.6b`** — drop a `.wav` (10-30 s of clean speech is plenty) into `/data/custom-voices/<anywhere>.wav`. Optionally drop a sibling `.txt` with a faithful transcript for higher clone fidelity. The voice appears in `GET /v1/audio/voices` on the next request — no restart. CUDA required.\n15. **Qwen3-TTS first synth is slow** — CUDA graph capture runs once after model load (~30-60 s). Subsequent synths are sub-second. If you're benchmarking, throw away the first call.\n16. **Qwen3-TTS ignores `speed`** — the model has no playback-rate control. Pass it for OpenAI compat; nothing happens. Only Kokoro honors `speed`.\n17. **Both TTS engines emit 24 kHz mono PCM** — Kokoro and Qwen3-TTS both output int16 24 kHz mono. ffmpeg re-encodes into your chosen `response_format`; `pcm` (raw, no container) hands back that same rate directly — check the `X-Sample-Rate` response header on Qwen3-TTS streaming responses if you need it confirmed per-request.\n18. **TTS `response_format=pcm` is for chaining** — raw int16 mono PCM, no container, no header. Use it when piping into another encoder or a real-time playback path. Otherwise stick with `mp3` (default) or `opus` for size.\n19. **TTS evicts loaded ASR and vice versa** — they share the same one-model-resident pool. Synthesizing with Kokoro after a transcribe burst incurs Kokoro's cold load. Same applies to Qwen3-TTS (plus the CUDA-graph capture re-runs on cold reload).\n\nFile v1.3.18:_meta.json\n\n{\n  \"ownerId\": \"kn79dhvmpjng4rp2jjk8k0v5xx80ccbk\",\n  \"slug\": \"talkies\",\n  \"version\": \"1.3.18\",\n  \"publishedAt\": 1787871031680\n}\n\nFile v1.3.18:references/setup.md\n\n# talkies setup\n\n## Requirements\n\n- Docker\n- `linux/amd64` host (no arm64 images — `nemo_toolkit[asr]` + chain doesn't resolve cleanly on aarch64)\n- Optional: NVIDIA GPU + NVIDIA Container Toolkit for the CUDA image (required for `qwen3-tts-0.6b` voice cloning)\n- ~3 GB disk for the CPU image, ~11 GB for the CUDA image\n- Additional disk for selected model weights; set `TALKIES_ENABLED_MODELS` to avoid downloading the full registry\n- ~4 GB RAM minimum (whisper-large-v3 needs the working set + overhead); 12 GB+ VRAM for the GPU-only models\n\n## Quick Install\n\n### CPU\n\nServes 2× Whisper + `canary-180m-flash` + `nemotron-3.5-asr-0.6b` (CPU-optimized, via parakeet.cpp), four selectable Sherpa Zipformer variants, and `vosk-small-en-us-0.15` for ASR, plus `kokoro-82m` and `kokoro-82m-nvidia` for TTS. The CUDA-only ASR models aren't worth running on CPU, and the Qwen3-TTS family is CUDA-only.\n\n```bash\ndocker run -d --name talkies \\\n  -v $HOME/talkies-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest\n```\n\n### CUDA\n\nServes all fourteen ASR models plus all three TTS engines / 4 backends (`kokoro-82m`, `kokoro-82m-nvidia`, the 5 Qwen3-TTS slugs, and `chatterbox-turbo`). Requires the NVIDIA Container Toolkit on the host.\n\nNemotron runs through the SHA-256-pinned upstream parakeet.cpp v0.5.0 CUDA 12\nbundle in this image, so both file transcription and native WebSocket sessions\nuse GPU offload. The matching CUDA 12.9 runtime libraries stay isolated under\n`/opt/parakeet` from the image's Python ML stack.\n\n```bash\ndocker run -d --name talkies \\\n  --gpus all \\\n  -v $HOME/talkies-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest-cuda\n```\n\nThe CUDA image expects `--gpus all`. Without a GPU assignment it retains its\n`TALKIES_DEVICE=cuda` image default, so model loading fails rather than silently\nfalling back. To use its CPU-compatible subset for debugging, explicitly set\n`-e TALKIES_DEVICE=cpu` and restrict `TALKIES_ENABLED_MODELS` to CPU-compatible\nslugs; GPU-only Qwen3-TTS slugs remain unavailable.\n\n**Verify:** `curl http://localhost:8000/healthz` returns `{\"ok\": true, \"device\": \"...\", \"models\": [...]}` once boot's done.\n\n**First boot:** the entrypoint downloads every enabled model into `/data/models/<slug>/` and creates `/data/files/` + `/data/custom-voices/`. Bind-mount `/data` so subsequent restarts are no-ops. Restrict the download set with `TALKIES_ENABLED_MODELS` to avoid pulling everything.\n\n## CPU vs CUDA Images\n\n| Image | Tag | Platforms | Models served | Image size |\n|---|---|---|---|---|\n| CPU | `psyb0t/talkies:latest` | `linux/amd64` | 2× Whisper, Canary-180m-Flash, Nemotron-3.5-ASR, Sherpa Zipformer ×4, Vosk, wav2vec2 + ZIPA phoneme, Kokoro-82M ×2 runtimes | ~3 GB |\n| CUDA | `psyb0t/talkies:latest-cuda` | `linux/amd64` | all fourteen ASR + Kokoro-82M ×2 runtimes + Qwen3-TTS ×5 + Chatterbox Turbo | ~11 GB |\n\nThe CPU image only ships ASR models that actually finish in a sane time without a GPU. Parakeet-TDT is autoregressive (slow on CPU). Canary-1B and Canary-Qwen-2.5B need the CUDA image with `--gpus all`; use the CPU image for CPU workloads. Kokoro-82M ships in both images — at 82M params it synthesizes faster than real-time on a 4-core CPU, no GPU needed. Chatterbox Turbo is CUDA-only for the same reason as the heavier ASR models: it runs on CPU but measures roughly 5-10x slower than real-time, so it is not registered as a CPU slug.\n\nBoth images bake `espeak-ng` into the runtime layer because Kokoro's G2P for es/fr/hi/it/pt routes through it via `misaki.espeak.EspeakG2P`. The Python `kokoro==0.9.4` package and its lightweight dependency chain (`misaki`, no `[ja]` / `[zh]` extras) are pinned alongside the rest of the ML stack in `Dockerfile` / `Dockerfile.cuda`.\n\nThe CUDA image additionally bakes the `faster-qwen3-tts==0.2.6` MIT wrapper and three builtin Qwen3 reference voices (`alloy`, `echo`, `fable`) under `/opt/talkies/qwen3-voices/`. The model weights (`Qwen/Qwen3-TTS-12Hz-0.6B-Base`, Apache-2.0) are downloaded into `/data/models/qwen3-tts-0.6b/` at first boot like every other model.\n\nThe CUDA image also bakes `chatterbox-tts==0.1.7` (MIT) and `s3tokenizer==0.3.0` (Apache-2.0) from a separate hash-pinned `requirements-chatterbox.txt`, installed `--no-deps` because their declared dependency metadata conflicts with the image's pinned torch/transformers and pulls tooling that has no place in a runtime image. The `ResembleAI/chatterbox-turbo` weights (MIT, ungated) land in `/data/models/chatterbox-turbo/` at first boot. Its voices come from `/data/custom-voices/` plus a `builtin` speaker shipped inside the checkpoint — the Qwen3 reference voices are deliberately not shared with it, since several are shorter than its 5-second reference-clip minimum.\n\n## Environment Variables\n\n### Auth + bind\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_AUTH_TOKEN` | (empty = no auth) | Bearer token required on every route except `/healthz`. Empty/unset = wide open (historical default — fine on private networks). When set, `Authorization: Bearer <token>` required on every HTTP request AND every MCP call. Compared with `hmac.compare_digest`. |\n\nContainer binds `0.0.0.0:8000` unconditionally. Control network exposure at `docker run` time:\n- `-p 127.0.0.1:8000:8000` — loopback-only on the host.\n- `-p 8000:8000` — all host interfaces.\n- For untrusted networks, combine the token with a reverse proxy doing TLS + rate limiting.\n\n### Device + model registry\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_DEVICE` | image default (`cpu` CPU / `cuda` CUDA) | `auto` picks `cuda` if available else `cpu`; it is an accepted override. Pin to a specific GPU with `cuda:N`. |\n| `TALKIES_MODELS_FILE` | `/app/models.json` | Path to the model registry JSON. Override to ship a custom subset. The CPU image copies `models-cpu.json` to this path; the CUDA image copies `models.json` here. |\n| `TALKIES_ENABLED_MODELS` | (empty = all from `models.json`) | Comma-separated slug whitelist. Restricts both the boot-time snapshot download and the queryable surface of `/v1/models`. Unknown slugs fail fast on startup. |\n| `TALKIES_PRELOAD` | (empty) | Comma-separated slugs to load into RAM/VRAM at boot, before uvicorn accepts requests. Skips cold-load on first transcription. Must be a subset of `TALKIES_ENABLED_MODELS`. |\n| `TALKIES_MODEL_MAX_CONCURRENCY` | `1` | Fallback number of simultaneous inference requests admitted per model across HTTP, MCP, WebSocket ASR, buffered TTS, and streaming TTS. Registry `max_concurrency` values take precedence. |\n| `TALKIES_MODEL_CONCURRENCY` | (empty) | Comma-separated `model-slug=limit` overrides, for example `nemotron-3.5-asr-0.6b=2,kokoro-82m=4`. Unknown, disabled, duplicate, malformed, or out-of-range entries fail at startup. |\n\nEach registry model may define `max_concurrency` from 1 through 1024. The\nbundled Nemotron entry defaults to two in both images. Only one model may own\nactive inference slots at a time, which prevents sibling model eviction while\na request is still using its backend.\n\n### Data dir\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_DATA_DIR` | `/data` | Base data dir. Model snapshots → `$TALKIES_DATA_DIR/models/<slug>/` (flat per-model dirs, no HF cache layout). Staged uploads + URL downloads → `$TALKIES_DATA_DIR/files/`. Qwen3-TTS custom clone voices → `$TALKIES_DATA_DIR/custom-voices/` (nested subdirs preserved as voice names). Bind-mount to persist across restarts. |\n\n**Security note on `$TALKIES_DATA_DIR/files/`:** staged uploads and cached URL downloads persist here **indefinitely** — nothing auto-expires them — and are enumerable by any caller via `GET /v1/files` (no per-caller isolation; see [Server-Side File Staging](../SKILL.md#server-side-file-staging-v1files) in SKILL.md). This is a shared bucket: an agent must only read/delete paths it staged itself, must never enumerate or delete other callers' files, and should clean up after its own workflow. Deploy with `TALKIES_AUTH_TOKEN` set by default, add per-caller isolation and retention limits at the deployment/proxy level if the deployment isn't fully trusted, and least-privilege network exposure otherwise.\n\n### Lifecycle (idle sweeper + load timeouts)\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_MODEL_TTL` | `600` (10 min) | Idle time before a loaded backend is unloaded by the sweeper. Bare number = seconds; also accepts Go-style `3h30m5s`, `45m`, `90s`. `0` disables auto-unload. |\n| `TALKIES_SWEEPER_INTERVAL` | `60` | How often the sweeper checks for idle models. |\n| `TALKIES_LOAD_TIMEOUT` | `300` | Parsed configuration reserved for a future model-load timeout; the current server does not apply it. |\n\n### Upload + download caps\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_MAX_UPLOAD_BYTES` | `104857600` (100 MB) | Reject `POST /v1/audio/transcriptions` multipart `file` and `PUT /v1/files/{path}` bodies larger than this with 413. |\n| `TALKIES_MAX_DOWNLOAD_BYTES` | `1073741824` (1 GiB) | Abort URL downloads (when `file_path` is an http(s) URL) larger than this. Larger default because downloads stream straight to disk, no in-memory buffering. |\n| `TALKIES_BLOCK_PRIVATE_DOWNLOADS` | `false` | Set to `true` to refuse URL downloads whose hostname resolves to private/loopback/link-local/multicast/reserved IPs. Default `false` because the typical self-hosted deployment is a LAN box fetching from another LAN box. Flip to `true` if exposed to untrusted clients. |\n\n### VAD knobs\n\nAudio longer than `TALKIES_VAD_CHUNK_THRESHOLD` seconds gets sliced through Silero VAD into ≤`TALKIES_VAD_MAX_SPEECH`-second speech regions before being handed to the backend.\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_VAD_CHUNK_THRESHOLD` | `30.0` | Audio longer than this (seconds) goes through VAD chunking. Shorter clips skip it. |\n| `TALKIES_VAD_MAX_SPEECH` | `28.0` | Max length of a single VAD-detected speech region (seconds). Should stay under Whisper's 30 s internal window. |\n| `TALKIES_VAD_MIN_SILENCE_MS` | `500` | Silero VAD param — minimum gap (ms) to consider a region break. |\n| `TALKIES_VAD_SPEECH_PAD_MS` | `200` | Silero VAD param — silence padding (ms) around each detected speech region. |\n| `TALKIES_VAD_THRESHOLD` | `0.5` | Silero VAD speech-probability threshold. Lower = more aggressive. |\n\n### Live ASR streaming\n\n`WS /v1/audio/transcriptions/stream` accepts headerless 16 kHz mono PCM16LE.\nIt is separate from the OpenAI-compatible upload route. See the repository's\n[`docs/streaming.md`](https://github.com/psyb0t/docker-talkies/blob/main/docs/streaming.md)\nfor the protocol and client examples.\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_STREAM_MAX_CONNECTIONS` | `4` | Maximum active ASR WebSockets per container. Streams may share one pinned model; attempts to switch models while one is active return a conflict. |\n| `TALKIES_STREAM_MAX_FRAME_BYTES` | `65536` | Maximum binary PCM frame size. Frames must be non-empty, contain whole 16-bit samples, and be 2–16777216 bytes. |\n| `TALKIES_STREAM_MAX_BUFFER_SECONDS` | `5` | Faster-whisper rolling-window budget. Must hold one configured maximum-size frame; native decoders process each frame directly. |\n| `TALKIES_STREAM_IDLE_TIMEOUT` | `30s` | Maximum wait between client messages before close code 4408. |\n| `TALKIES_STREAM_MAX_DURATION` | `4h` | Maximum accepted audio duration per WebSocket. |\n\n### Qwen3-TTS streaming\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_QWEN3_STREAM_CHUNK_SIZE` | `8` | Codec steps decoded per yielded chunk when `response_format=pcm` streams from a `qwen3_tts` backend (~1 s of audio per 12 steps). Only relevant to that streaming path. |\n\n### Logging\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_LOG_LEVEL` (falls back to `LOG_LEVEL`) | `info` | `debug` / `info` / `warn` / `error` / `fatal` (case-insensitive; `warning` / `critical` also accepted). Unrecognized values fail fast at startup. JSON structured logs on stdout. **`debug` logs full request/response bodies** (TTS input text, cloned-voice reference transcripts, ASR transcripts) — PII; a one-time WARNING fires at startup when active. |\n\n### Internal\n\n| Var | Default | What it does |\n|---|---|---|\n| `HF_HUB_OFFLINE` | `1` (in image) | Refuse network calls from HuggingFace Hub at runtime. The entrypoint transparently unsets it for the one-shot prefetch step so the initial download works; the server process itself runs offline. Don't touch unless debugging. |\n\n## Common Configurations\n\n```bash\n# Restrict to just the small/fast models (saves first-boot download time).\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_ENABLED_MODELS=whisper-large-v3-turbo,canary-180m-flash \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Preload at boot so the first request doesn't pay the cold-load tax.\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_ENABLED_MODELS=whisper-large-v3-turbo \\\n  -e TALKIES_PRELOAD=whisper-large-v3-turbo \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Bearer auth on a public-facing deployment.\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_AUTH_TOKEN=$(openssl rand -hex 32) \\\n  -e TALKIES_BLOCK_PRIVATE_DOWNLOADS=true \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Loopback only (rely on reverse proxy for external access).\ndocker run -d -p 127.0.0.1:8000:8000 \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Disable auto-unload (keep model resident forever).\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_MODEL_TTL=0 \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Bump upload + download caps for huge files.\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_MAX_UPLOAD_BYTES=1073741824 \\\n  -e TALKIES_MAX_DOWNLOAD_BYTES=10737418240 \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Pin to a specific GPU on a multi-GPU host.\ndocker run -d --gpus '\"device=1\"' -p 8000:8000 \\\n  -e TALKIES_DEVICE=cuda:0 \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest-cuda\n```\n\n## Ports\n\n| Port | Service |\n| ---- | ------- |\n| 8000 | HTTP API + MCP (`/v1/mcp`) on the same port |\n\nContainer binds `0.0.0.0:8000` unconditionally — there are no `TALKIES_HOST` / `TALKIES_PORT` env vars (they were removed in v0.2.0). Use `-p` at `docker run` time for whatever host port mapping you want.\n\n## Customizing the Model Registry\n\nThe image ships with `models.json` (CUDA) or `models-cpu.json` (CPU) baked in. Override without rebuilding by bind-mounting your own:\n\n```bash\ndocker run -d --name talkies \\\n  -v $HOME/talkies-data:/data \\\n  -v $PWD/my-models.json:/app/models.json:ro \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest\n```\n\nOr point `TALKIES_MODELS_FILE` at a different path inside the container.\n\nFile structure:\n\n```json\n{\n  \"models\": {\n    \"your-asr-slug\": {\n      \"repo\": \"huggingface-org/repo-name\",\n      \"executor\": \"whisper\",\n      \"default_source_lang\": \"en\",\n      \"default_target_lang\": \"en\",\n      \"default_task\": \"asr\",\n      \"languages\": [\"en\"]\n    },\n    \"your-tts-slug\": {\n      \"repo\": \"huggingface-org/tts-repo-name\",\n      \"executor\": \"kokoro\",\n      \"modality\": \"tts\",\n      \"default_voice\": \"af_heart\",\n      \"languages\": [\"en\"]\n    }\n  }\n}\n```\n\n| Field | Required | Notes |\n|---|---|---|\n| `repo` | yes | HuggingFace repo id. Pulled via `snapshot_download(local_dir=$TALKIES_DATA_DIR/models/<slug>)` — flat directory keyed by slug, no HF cache indirection. |\n| `revision` | no | Immutable Hugging Face commit SHA to download. Pin this for reproducible custom registries. |\n| `executor` | yes | One of `whisper`, `parakeet`, `parakeet_cpp`, `canary_multitask`, `canary_salm`, `sherpa`, `sherpa_offline_ctc`, `vosk`, `kokoro`, `kokoro_nvidia`, `qwen3_tts`, `chatterbox`, `wav2vec2_phoneme`. Other values fail startup — the allowlist is `VALID_EXECUTORS` in `src/talkies/config.py`, and `load_registry()` runs at server import, so an unknown executor stops the whole process rather than disabling one model. |\n| `modality` | no | `asr` (default) or `tts`. Drives endpoint guards (`/v1/audio/transcriptions` requires ASR; `/v1/audio/speech` requires TTS) and the `modality` field on `/v1/models` entries. The `kokoro`, `qwen3_tts` and `chatterbox` executors imply `tts`; the seven ASR executors imply `asr`. |\n| `download_patterns` | no | Non-empty list of static repository-relative paths passed to Hugging Face `snapshot_download(..., allow_patterns=...)`. Use it to limit a multi-variant repository to the files selected by this registry entry. |\n| `default_source_lang` | no | ASR only. Used when the request omits `language`. |\n| `default_target_lang` | no | ASR only. Used by Canary multitask for translation tasks. |\n| `default_task` | no | ASR only. `asr` (transcribe) or `s2t_translation` (Canary multitask only). Default `asr`. |\n| `default_voice` | no | TTS only. Used when the request omits `voice`. Falls back to the first voice the backend reports. For `qwen3_tts`, the voice name is a path relative to the voices dir (`alloy`, `team-a/jane`). |\n| `default_language` | no | `qwen3_tts` only. Default reference-clip language label (defaults to `English`). Overridden per-voice by a sibling `.lang` file next to the wav. |\n| `languages` | no | Informational only — listed in error messages, not enforced. |\n| `dependencies` | no | List of extra HuggingFace repo ids the executor needs at load time (e.g. `canary-qwen-2.5b` instantiates a Qwen3 tokenizer separately). Each is `snapshot_download`'d into the standard HF cache (`HF_HOME`) at entrypoint. |\n\n### Common customization: translation slugs\n\nThe shipped `models.json` ships every Canary slug with `default_task=asr`, so out of the box the API only transcribes. To enable translation (Canary-1B-Flash covers en↔de/fr/es), add a translation-specific slug:\n\n```json\n{\n  \"models\": {\n    \"canary-1b-flash-de2en\": {\n      \"repo\": \"nvidia/canary-1b-flash\",\n      \"executor\": \"canary_multitask\",\n      \"default_source_lang\": \"de\",\n      \"default_target_lang\": \"en\",\n      \"default_task\": \"s2t_translation\",\n      \"languages\": [\"de\"]\n    },\n    \"canary-1b-flash-en2de\": {\n      \"repo\": \"nvidia/canary-1b-flash\",\n      \"executor\": \"canary_multitask\",\n      \"default_source_lang\": \"en\",\n      \"default_target_lang\": \"de\",\n      \"default_task\": \"s2t_translation\",\n      \"languages\": [\"en\"]\n    }\n  }\n}\n```\n\nMultiple slugs can point at the same HF repo — talkies loads the underlying weights once and changes the prompt format per slug.\n\n### Common customization: restricting to one model\n\nFor a single-purpose deployment, ship a one-entry registry to skip pulling everything:\n\n```json\n{\n  \"models\": {\n    \"whisper-large-v3-turbo\": {\n      \"repo\": \"deepdml/faster-whisper-large-v3-turbo-ct2\",\n      \"executor\": \"whisper\",\n      \"default_source_lang\": \"en\",\n      \"languages\": [\"en\"]\n    }\n  }\n}\n```\n\nEquivalent to setting `TALKIES_ENABLED_MODELS=whisper-large-v3-turbo` against the default registry — but with a custom registry you can add slugs that aren't in the shipped one.\n\n## Qwen3-TTS Custom Voices\n\n**Acceptable use:** only supply reference voice samples you're authorized to process, and only with the speaker's informed consent. Voice cloning reproduces someone's actual timbre/prosody — never use it to impersonate a real person without consent, or for fraud, deception, or any form of unauthorized voice replication.\n\n`qwen3-tts-0.6b` is a voice-cloning TTS — it takes a reference `.wav` and clones the speaker's timbre / prosody onto whatever text you supply. The voice catalog is built from two on-disk dirs that are merged at request time (live, no restart):\n\n| Dir | Where it lives | Origin tag | Purpose |\n|---|---|---|---|\n| Builtin | `/opt/talkies/qwen3-voices/` (baked into the CUDA image) | `builtin` | Three curated samples (`alloy`, `echo`, `fable`) so the model works out of the box. |\n| Custom | `/data/custom-voices/` (host-mounted) | `custom` | Your reference clips. Drop in, get back. |\n\nVoice names are the wav's path relative to the parent dir with `.wav` stripped. Nested subdirs are preserved:\n\n```\n$HOME/talkies-data/custom-voices/\n├── jane.wav              → voice \"jane\"\n├── jane.txt              # optional reference transcript\n├── jane.lang             # optional language label (defaults \"English\")\n└── team-a/\n    └── narrator-bob.wav  → voice \"team-a/narrator-bob\"\n```\n\nCustom voices **shadow** builtin voices with the same name — dropping `custom-voices/alloy.wav` overrides the builtin `alloy` (its `origin` field on `/v1/audio/voices` flips from `builtin` to `custom`).\n\n**Sibling metadata** next to each `<name>.wav`:\n- `<name>.txt` — reference transcript for the clip. Optional; the model accepts an empty string. Clone fidelity is noticeably better with a faithful transcript.\n- `<name>.lang` — language label string passed through to the model. Optional; defaults to `English`. Use this for non-English reference clips.\n\n**Recommended reference clips:**\n- 10-30 s of clean speech from the target speaker.\n- No background music, no overlapping voices, low noise floor.\n- 16+ kHz, mono preferred (model resamples internally but garbage-in-garbage-out applies).\n\n**Use a custom voice:**\n\n```bash\nmkdir -p $HOME/talkies-data/custom-voices/team-a\ncp jane-reading.wav $HOME/talkies-data/custom-voices/team-a/jane.wav\necho \"And the silken sad uncertain rustling of each purple curtain.\" \\\n  > $HOME/talkies-data/custom-voices/team-a/jane.txt\n\ncurl -s http://localhost:8000/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"qwen3-tts-0.6b\",\n        \"voice\": \"team-a/jane\",\n        \"input\": \"Hello from a cloned voice.\",\n        \"response_format\": \"wav\"\n      }' \\\n  --output cloned.wav\n```\n\n**Path-traversal guard:** symlinks under `custom-voices/` whose `resolve()` escapes the dir are skipped at scan time, so a hostile mount can't be used to read arbitrary host files as a voice prompt. Symlinks pointing back into the same dir are fine.\n\n**CUDA only.** Qwen3-TTS hard-fails on CPU (`FasterQwen3TTS.from_pretrained` raises `ValueError`). The model surfaces as `loaded: false` until the first request; first-request load includes CUDA-graph capture (~30-60 s on a mid-range GPU). Subsequent generations are sub-second.\n\n## OpenClaw / ClawHub Config\n\n```bash\nexport TALKIES_URL=http://localhost:8000\nexport TALKIES_AUTH_TOKEN=<token>  # only if the server requires it\n```\n\nOr via `~/.openclaw/openclaw.json`:\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"talkies\": {\n        \"env\": {\n          \"TALKIES_URL\": \"http://localhost:8000\",\n          \"TALKIES_AUTH_TOKEN\": \"<token>\"\n        }\n      }\n    }\n  }\n}\n```\n\n## Management\n\n```bash\ndocker logs -f talkies    # tail logs\ndocker stop talkies       # stop\ndocker rm talkies         # remove\ndocker pull psyb0t/talkies:latest  # update\n```\n\nWatch what's loaded right now:\n\n```bash\ncurl -s http://localhost:8000/api/ps | jq\n```\n\nFree memory between jobs:\n\n```bash\ncurl -s -X POST http://localhost:8000/unload | jq\n```\n\n## Logs\n\n`docker logs talkies` covers everything. Look for:\n\n- `entrypoint:` lines on boot — model snapshot downloads, device detection.\n- `INFO talkies.server` lines on each request — model load events, transcribe timings.\n- `WARNING` / `ERROR` lines for backend failures.\n\nAt `info` (default) and above, the server does not log auth tokens, request/response bodies, or audio bytes — it logs the model slug, request id, duration, and result size.\n\n**At `debug`, this changes: full request/response content is logged**, including TTS input text + `instructions`, cloned-voice reference transcripts, and ASR transcripts (`src/talkies/server.py` request/response `log.debug(...)` calls, gated behind `log.isEnabledFor(logging.DEBUG)`). This is PII. A one-time WARNING fires at startup when `TALKIES_LOG_LEVEL=debug` is active (`src/talkies/logging.py`). **Never run `debug` level in production against real user data** — use it only for local troubleshooting with synthetic/throwaway input.\n\n## Public Access via Reverse Proxy (optional)\n\ntalkies binds `0.0.0.0:8000` inside the container. For public exposure, terminate TLS at a reverse proxy (Caddy / Traefik / nginx) and combine with `TALKIES_AUTH_TOKEN`.\n\nCaddy example:\n\n```caddy\ntalkies.example.com {\n    reverse_proxy localhost:8000\n}\n```\n\nSet the auth token on the talkies container so even if Caddy is misconfigured, the upstream still requires `Authorization: Bearer`. Don't rely on the proxy alone.\n\nFor Cloudflare Tunnel / Tailscale, the same logic applies — the tunnel provides transport security, the bearer token provides app-layer auth.\n\nFile v1.3.18:skill-card.md\n\n## Description:\n\ntalkies helps agents use a self-hosted OpenAI-compatible speech service for transcription, live ASR, speech synthesis, subtitles, server-side file staging, and MCP-based ASR workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[psyb0t](https://clawhub.ai/user/psyb0t)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agents use talkies to run speech-to-text, text-to-speech, subtitle generation, live transcription, and batch transcription workflows against a trusted self-hosted Talkies server. It is useful when an OpenAI-compatible audio API, optional MCP ASR tools, and local control over speech models are required.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Speech input, synthesized text, and voice-cloning reference samples are sent to the configured Talkies server.\n\nMitigation: Point TALKIES_URL only at an instance the operator runs or explicitly trusts, prefer HTTPS outside localhost or a protected LAN, and avoid sending sensitive media or text to untrusted servers.\n\nRisk: An unconfigured Talkies deployment can expose the API, staged files, and management endpoints to anyone who can reach the port.\n\nMitigation: Bind the service to localhost or a protected network, set TALKIES_AUTH_TOKEN for shared or public deployments, and use a reverse proxy with TLS and rate limiting when exposing it externally.\n\nRisk: Server-side URL file_path fetching can access operator-side network resources and persist downloaded media on the server.\n\nMitigation: Enable TALKIES_BLOCK_PRIVATE_DOWNLOADS before accepting untrusted callers, pass only trusted URLs, and delete cached downloads after the workflow is complete.\n\nRisk: Server-side staged files are persistent, shared, and enumerable by callers with API access.\n\nMitigation: Do not stage sensitive media on shared instances without authentication, only read or delete files created for the current workflow, and clean up staged files promptly.\n\nRisk: Voice cloning and expressive TTS can be misused to impersonate real people.\n\nMitigation: Clone or synthesize a real person's voice only with explicit authorization and consent, and remove custom voice samples when no longer needed.\n\nRisk: Setup and workflow examples execute local docker, curl, and ffmpeg commands, and latest container tags may change over time.\n\nMitigation: Review commands before execution, run them in an appropriate environment, and prefer pinned image digests for repeatable deployments.\n\nRisk: Debug logging can expose request and response content, including transcripts, TTS text, and voice reference transcripts.\n\nMitigation: Keep production deployments at info or higher log levels and use debug logging only for local troubleshooting with synthetic or disposable data.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/psyb0t/skills/talkies)\n- [talkies setup guide](references/setup.md)\n- [Talkies repository](https://github.com/psyb0t/docker-talkies)\n- [Streaming protocol documentation](https://github.com/psyb0t/docker-talkies/blob/main/docs/streaming.md)\n- [Model Context Protocol](https://modelcontextprotocol.io)\n- [speaches project](https://github.com/speaches-ai/speaches)\n\n## Skill Output:\n\n**Output Type(s):** [text, shell commands, configuration, guidance, files]\n\n**Output Format:** [Markdown guidance with shell commands, API examples, JSON or plain-text transcripts, SRT/VTT subtitles, and generated audio files.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs depend on the configured TALKIES_URL, enabled ASR/TTS models, response_format, authentication, and staged-file lifecycle.]\n\n## Skill Version(s):\n\n1.3.18 (source: ClawHub release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.3.17: 5 files, 31071 bytes\n\nFiles: references/setup.md (24535b), scripts/bulk_transcribe.sh (3497b), skill-card.md (2550b), SKILL.md (48735b), _meta.json (127b)\n\nFile v1.3.17:SKILL.md\n\n---\nname: talkies\ndescription: Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 12 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.\nhomepage: https://github.com/psyb0t/docker-talkies\nuser-invocable: true\npermissions:\n  network: \"Outbound HTTP to the configured TALKIES_URL; the Talkies server also fetches URLs supplied as file_path.\"\n  shell: \"Documented setup and workflow examples invoke local curl, ffmpeg, and docker commands.\"\n  filesystem: \"Reads and writes server-side staged files through /v1/files; the skill itself does not access the local filesystem.\"\nmetadata:\n  { \"openclaw\": { \"emoji\": \"🎙️\", \"primaryEnv\": \"TALKIES_URL\", \"requires\": { \"bins\": [\"docker\", \"curl\"] } } }\n---\n\n# talkies\n\nSelf-hosted speech service — ASR and TTS, one container. OpenAI-compatible wire shape on both endpoints; point an OpenAI client at it, change the model slug, done.\n\nASR (`POST /v1/audio/transcriptions`): twelve bundled slugs — `whisper-large-v3`, `whisper-large-v3-turbo`, `parakeet-tdt-0.6b-v3`, `nemotron-3.5-asr-0.6b`, `canary-180m-flash`, `canary-1b-flash`, `canary-qwen-2.5b`, four selectable English Sherpa Zipformer variants, and `vosk-small-en-us-0.15`.\n\nTTS (`POST /v1/audio/speech`): 3 engines / 4 backends across 8 slugs — `kokoro-82m` (PyTorch) and `kokoro-82m-nvidia` (ONNX/ORT) with 41 baked voices across en/es/fr/hi/it/pt, plus the CUDA-only Qwen3-TTS family: `qwen3-tts-0.6b` / `qwen3-tts-1.7b` (voice cloning from reference clips), `qwen3-tts-0.6b-custom` / `qwen3-tts-1.7b-custom` (9 preset speakers), `qwen3-tts-1.7b-design` (voice from an NL description), plus the CUDA-only `chatterbox-turbo` (English only; 19 inline emotion tags; clones from a reference `.wav` with no transcript). Discover voices via `GET /v1/audio/voices`.\n\nExtras: live PCM ASR over WebSocket, stereo diarization on transcription, URL `file_path` fetching, server-side file staging, MCP endpoint with 6 ASR-side tools, optional bearer-token auth.\n\nFor installation, configuration, and container setup, see [references/setup.md](references/setup.md).\n\n## Security & safety\n\nThis skill is **not** low-risk to orchestrate blindly — it issues local shell commands (`curl`, `ffmpeg`, `docker` in the typical deployment/setup path) and outbound HTTP requests to whatever `$TALKIES_URL` points at:\n\n- **Outbound HTTP to an operator-chosen host — data leaves your host.** Every command in this skill, including TTS `input` text and voice-cloning reference samples, is sent to whatever `$TALKIES_URL` points at — a server the skill does not control or vet. Point it only at an instance you run or explicitly trust; prefer HTTPS.\n- **`file_path` URL fetches happen server-side, not client-side.** Passing a URL causes the *talkies server* to download it — see [URL `file_path` (Download + Cache)](#url-file_path-download--cache) below.\n- **Staged files persist server-side until explicitly deleted**, and are enumerable by anyone who can reach the API — see [Server-Side File Staging](#server-side-file-staging-v1files) below.\n- **Shared file staging** — staged files (`/v1/files`) live in one namespace with no per-caller ownership: any caller can see or remove any staged file. Only remove paths you staged in the current task, and treat file management as admin-only on a shared instance — see [Server-Side File Staging](#server-side-file-staging-v1files) below.\n- **No auth by default** — `TALKIES_AUTH_TOKEN` is opt-in; an unconfigured server is wide open to anyone who can reach the port. Require the token and least-privilege network exposure in any non-trusted environment.\n- **Voice cloning requires consent.** Only clone or synthesize a voice you have explicit authorization/consent to use — synthesized speech of a real person can be used for impersonation, fraud, or deception. See [Qwen3-TTS Custom Voices](references/setup.md#qwen3-tts-custom-voices).\n- **Local shell execution.** `references/setup.md` and workflow examples run `docker`, `curl`, and `ffmpeg` directly on the local host — review commands before running them against unfamiliar hosts/images.\n\n## When To Use\n\n- Transcribe audio files (any format ffmpeg decodes — WAV, MP3, M4A, FLAC, OGG, WebM, Opus, MP4 audio).\n- Generate SRT/VTT subtitles for video.\n- Transcribe podcasts, lectures, interviews, voicemails, calls.\n- Stream 16 kHz mono PCM from a microphone or decoded live-audio source through `WS /v1/audio/transcriptions/stream`.\n- Stereo two-mic recordings → per-speaker diarized output (`L:` / `R:` channel tagging).\n- German/French/Spanish ↔ English speech-to-text translation via Canary-1B-Flash.\n- Synthesize speech from text via Kokoro-82M — English (American + British), Spanish, French, Hindi, Italian, Portuguese.\n- Voice-clone speech via Qwen3-TTS-0.6B from a reference `.wav` you provide — drop into `/data/custom-voices/`, immediately appears under `GET /v1/audio/voices` with `origin=custom`. **Only clone voices you're authorized to use, with the speaker's consent** — see [Qwen3-TTS Custom Voices](references/setup.md#qwen3-tts-custom-voices).\n- Voice-clone via `chatterbox-turbo` from the same `/data/custom-voices/` drop — no sibling transcript needed, but the clip must be longer than 5 seconds. Same consent requirement applies.\n- Emotive English delivery via `chatterbox-turbo` — write tags inline in `input`, e.g. `Oh no [sigh] not again.` The tokenizer defines exactly 19: `[angry]` `[fear]` `[surprised]` `[whispering]` `[advertisement]` `[dramatic]` `[narration]` `[crying]` `[happy]` `[sarcastic]` `[clear throat]` `[sigh]` `[shush]` `[cough]` `[groan]` `[sniff]` `[gasp]` `[chuckle]` `[laugh]`. Anything else is read as literal text.\n- Drop-in replacement for `api.openai.com/v1/audio/transcriptions` and `api.openai.com/v1/audio/speech` in existing client code.\n\n## When NOT To Use\n\n- OpenAI-compatible live ASR — `/v1/audio/transcriptions` remains request/response only. Use talkies-specific `WS /v1/audio/transcriptions/stream` for raw PCM. (TTS has one streaming exception: `qwen3-tts-*` + `response_format=pcm` streams chunked PCM — see [Streaming PCM (Qwen3-TTS)](#streaming-pcm-qwen3-tts).)\n- Speaker identification from voice (only stereo-channel diarization is supported, not voice clustering).\n- Per-request `prompt` / `temperature` on `/v1/audio/transcriptions` — accepted for OpenAI compat, **ignored**. (`instructions` on `/v1/audio/speech` is different: Kokoro ignores it, but Qwen3-TTS honors it in most modes — see [Request Body](#request-body).)\n- Japanese / Chinese TTS — Kokoro upstream supports them but talkies filters those voices out (they need the `misaki[ja]` / `misaki[zh]` extras).\n- Kokoro on OpenAI aliases (`alloy`, `echo`, `fable`, `onyx`, `nova`, `shimmer`) — Kokoro exposes its native voice names only (`af_*`, `bm_*`, etc.). Map client-side. (Qwen3-TTS does ship `alloy` / `echo` / `fable` as builtin voice slugs, but they're voice-cloned samples, not OpenAI's voices — there's no audio compatibility.)\n- `qwen3-tts-0.6b` on CPU — voice cloning hard-fails without CUDA at load time. The `faster_qwen3_tts` upstream raises `ValueError` on non-CUDA devices; talkies surfaces this as a load failure on the first request.\n- `qwen3-tts-0.6b` `speed` parameter — Qwen3-TTS has no playback-rate control. Field is accepted for OpenAI compat but **ignored** (only Kokoro honors `speed`; `chatterbox-turbo` ignores it too).\n- `chatterbox-turbo` for anything but English — it is an English-only checkpoint. Use Kokoro or Qwen3-TTS for other languages.\n- `chatterbox-turbo` on CPU — the slug is registered in the CUDA image only. The model does run on CPU but measures roughly 5-10x slower than realtime, so it is not offered as a CPU slug.\n- `chatterbox-turbo` with default settings where output must be unwatermarked. Every waveform carries Resemble AI's PerTh neural watermark, which the upstream package applies unconditionally. The server can turn it off: set `TALKIES_CHATTERBOX_WATERMARK=false` on the container. That is an operator-side setting, not a per-request one, so a caller cannot change it through the API.\n- `chatterbox-turbo` reference clips of 5 seconds or shorter — rejected with a 400. Supply a longer clip.\n- arm64 hosts — `linux/amd64` only.\n\n## Setup\n\nThe container should already be running. Set the base URL:\n\n```bash\nexport TALKIES_URL=http://localhost:8000\n```\n\nIf the server has `TALKIES_AUTH_TOKEN` set, export it too:\n\n```bash\nexport TALKIES_AUTH_TOKEN=<your-token>\n# every request below needs: -H \"Authorization: Bearer $TALKIES_AUTH_TOKEN\"\n```\n\n**Verify:** `curl $TALKIES_URL/healthz` returns `{\"ok\": true, \"device\": \"...\", \"models\": [...]}`.\n\nFor install / configuration / env vars / CPU vs CUDA images / custom model registry, see [references/setup.md](references/setup.md).\n\n## Quick Start\n\n```bash\n# Discover what's available.\ncurl -s $TALKIES_URL/v1/models | jq\n\n# Simplest transcribe — file upload, JSON response.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq\n\n# Same call, but the audio lives at a URL — talkies downloads + caches it.\n# WARNING: the URL is fetched by the talkies SERVER, not this client, and the\n# download is cached persistently on disk (see \"URL file_path\" below). Don't\n# pass URLs to private/sensitive media unless you trust the server operator.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq\n\n# Full Whisper-shape JSON with per-segment + per-word timestamps.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"response_format=verbose_json\" | jq\n\n# SRT subtitles.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@lecture.mp3\" \\\n  -F \"model=whisper-large-v3\" \\\n  -F \"response_format=srt\" > lecture.srt\n\n# Discover TTS voices, then synthesize an MP3.\ncurl -s $TALKIES_URL/v1/audio/voices | jq\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"kokoro-82m\",\n        \"input\": \"Hello from talkies.\",\n        \"voice\": \"af_heart\",\n        \"response_format\": \"mp3\"\n      }' \\\n  --output hello.mp3\n```\n\n## Live streaming ASR\n\n`WS /v1/audio/transcriptions/stream` is talkies-specific rather than an OpenAI\nendpoint. It accepts headerless 16 kHz, mono, signed 16-bit little-endian PCM.\nSend one JSON `start` message, wait for `ready`, stream binary PCM frames, then\nsend `{\"type\":\"end\"}` for `final` and `stats` (or `{\"type\":\"cancel\"}` to\ndiscard the stream). Use a configured ASR slug with streaming support; native\nsessions are available through parakeet.cpp, Sherpa-ONNX, and Vosk, while\nfaster-whisper uses a bounded rolling window. When authentication is enabled,\nsend the bearer token in the WebSocket upgrade header, never in the URL.\n\nSee the repository's [`docs/streaming.md`](https://github.com/psyb0t/docker-talkies/blob/main/docs/streaming.md)\nfor the exact start/event shapes, a Python client, limits, and custom Sherpa or\nVosk model-registry entries.\n\n## Supported Models\n\n### ASR\n\n| Slug | Family | CPU | CUDA | Languages | Strength |\n|---|---|---|---|---|---|\n| `whisper-large-v3` | faster-whisper | yes | yes | 99 auto-detect | best accuracy, slowest |\n| `whisper-large-v3-turbo` | faster-whisper | yes | yes | 99 auto-detect | sweet spot — fast, accurate |\n| `parakeet-tdt-0.6b-v3` | NeMo TDT | no | yes | English only | very fast on GPU |\n| `nemotron-3.5-asr-0.6b` | parakeet.cpp / ggml | yes | yes | 40+ locales, auto-detect | CPU or CUDA multilingual; pin via `language=` |\n| `canary-180m-flash` | NeMo Canary | yes | yes | English only (small) | smallest, runs anywhere |\n| `canary-1b-flash` | NeMo Canary | no | yes | en/de/fr/es + translation | multilingual, translation |\n| `canary-qwen-2.5b` | NeMo SALM | no | yes | English only | best English accuracy (no timestamps) |\n| `sherpa-zipformer-en-left-64` | Sherpa-ONNX Zipformer | yes | yes | English only | native live ASR, lower left context |\n| `sherpa-zipformer-en-left-128` | Sherpa-ONNX Zipformer | yes | yes | English only | native live ASR, higher left context |\n| `sherpa-zipformer-en-int8-left-64` | Sherpa-ONNX Zipformer INT8 | yes | yes | English only | smaller native live ASR variant |\n| `sherpa-zipformer-en-int8-left-128` | Sherpa-ONNX Zipformer INT8 | yes | yes | English only | smaller native live ASR variant |\n| `vosk-small-en-us-0.15` | Vosk | yes | yes | English only | native live ASR; CPU decoder |\n\nPick by use case:\n- **General-purpose:** `whisper-large-v3-turbo`.\n- **English-only, max accuracy on GPU:** `canary-qwen-2.5b` (but no per-segment timestamps).\n- **Translation EN↔DE/FR/ES:** `canary-1b-flash` (requires custom model registry — see [Translation](#translation)).\n- **Low-overhead live English ASR:** one of the Sherpa Zipformer variants or `vosk-small-en-us-0.15`; Sherpa uses CUDA in the CUDA image, while Vosk remains CPU-decoded.\n\n### TTS\n\n3 engines / 4 backends across 8 slugs. Kokoro ships in two runtimes (`kokoro-82m` PyTorch, `kokoro-82m-nvidia` ONNX/ORT) — same weights, same voice catalog, same wire format. Qwen3-TTS ships 5 CUDA-only slugs across three modes (base cloning / custom_voice preset speakers / voice_design). Mode is implicit in the slug — see [Qwen3-TTS Modes](#qwen3-tts-modes). `chatterbox-turbo` (CUDA-only, English) rounds out the set — 19 inline emotion tags, transcript-free voice cloning; see [When To Use](#when-to-use) above.\n\n| Slug | Family | Mode | CPU | CUDA | Languages | Voices |\n|---|---|---|---|---|---|---|\n| `kokoro-82m` | Kokoro (PyTorch in-process, 24 kHz) | — | yes | yes | en (US + UK), es, fr, hi, it, pt | 41 baked (discover via `GET /v1/audio/voices`) |\n| `kokoro-82m-nvidia` | Kokoro (ONNX via ORT, 24 kHz) | — | yes | yes | en (US + UK), es, fr, hi, it, pt | 41 baked (same catalog as `kokoro-82m`) |\n| `qwen3-tts-0.6b` | Qwen3-TTS (24 kHz) | base | no | yes | 17 (en, zh, ja, ko, fr, de, es, it, pt, ru, vi, th, id, ar, tr, pl, nl) | 3 builtin samples + any `.wav` under `/data/custom-voices/` |\n| `qwen3-tts-1.7b` | Qwen3-TTS (24 kHz) | base | no | yes | 10 (en, zh, ja, ko, fr, de, es, it, pt, ru) | 3 builtin samples + any `.wav` under `/data/custom-voices/` |\n| `qwen3-tts-0.6b-custom` | Qwen3-TTS (24 kHz) | custom_voice | no | yes | en, zh, ja, ko | 9 preset speakers (`instructions` dropped — 0.6B limitation) |\n| `qwen3-tts-1.7b-custom` | Qwen3-TTS (24 kHz) | custom_voice | no | yes | en, zh, ja, ko | 9 preset speakers + emotion via `instructions` |\n| `qwen3-tts-1.7b-design` | Qwen3-TTS (24 kHz) | voice_design | no | yes | en, zh, ja, ko | voice synthesized from NL description in `instructions` (required) |\n| `chatterbox-turbo` | Chatterbox Turbo (24 kHz, buffered only) | — | no | yes | English only | `builtin` speaker, or any `.wav` (>5s) under `/data/custom-voices/` |\n\nPick by use case:\n- **General-purpose multi-voice TTS:** `kokoro-82m` — fast, 41 baked voices, runs on CPU. Use `kokoro-82m-nvidia` for the ONNX/ORT execution path (CUDA EP on the CUDA image, CPU EP otherwise).\n- **Voice cloning from a reference clip:** `qwen3-tts-0.6b` / `qwen3-tts-1.7b` — drop a `.wav` into `/data/custom-voices/`, immediately usable. CUDA required.\n- **Preset speakers (no reference WAV):** `qwen3-tts-0.6b-custom` / `qwen3-tts-1.7b-custom` — 9 baked speakers; the 1.7B honours `instructions` for emotion. CUDA required.\n- **Invent a voice from a description:** `qwen3-tts-1.7b-design` — the NL description goes in `instructions`. CUDA required.\n- **Expressive English delivery, transcript-free cloning:** `chatterbox-turbo` — 19 inline emotion tags in `input`, clones from a bare `.wav` (no sibling transcript needed, clip must be longer than 5 seconds). English only, CUDA required, output always carries a neural watermark.\n\n`canary-qwen-2.5b` produces no segment/word timestamps — `verbose_json.segments` and `.words` come back empty, `srt`/`vtt` collapse to a single full-duration cue. Transcription itself is whole-file. Use a Whisper or Canary multitask slug if you need timing.\n\n## API — `POST /v1/audio/transcriptions`\n\nMultipart form. Same field names as OpenAI's transcription endpoint where they overlap.\n\n### Request Fields\n\n| Field | Required | Default | Notes |\n|---|---|---|---|\n| `file` | one of `file`/`file_path` | — | Audio file. Capped at `TALKIES_MAX_UPLOAD_BYTES` (default 100 MB). |\n| `file_path` | one of `file`/`file_path` | — | Either a path under the staging area (`/v1/files`) or an `http(s)://` URL (downloaded + cached server-side). Not subject to the 100 MB upload cap; URL downloads capped by `TALKIES_MAX_DOWNLOAD_BYTES` (default 1 GiB). |\n| `model` | yes | — | One of the configured slugs (see `GET /v1/models`). Unknown → 404. |\n| `language` | no | model default | ISO-639-1 code. Whisper auto-detects when omitted; Canary uses its `default_source_lang`. |\n| `response_format` | no | `json` | `json` / `text` / `verbose_json` / `srt` / `vtt`. |\n| `timestamp_granularities[]` | no | — | Accepted for OpenAI compat; ignored — `verbose_json` always emits both segment + word. |\n| `prompt` | no | — | **Accepted, ignored.** |\n| `temperature` | no | — | **Accepted, ignored.** |\n| `diarization` | no | `false` | Stereo-channel diarization. Requires 2-channel input — mono returns 400. |\n\nExactly one of `file` or `file_path` must be set — passing both or neither returns 400.\n\n### Response Formats\n\n| `response_format` | Content-Type | Shape |\n|---|---|---|\n| `json` (default) | `application/json` | `{\"text\": \"...\"}` — just the transcript. |\n| `text` | `text/plain` | The transcript as plain text. |\n| `verbose_json` | `application/json` | Full Whisper shape — `task`, `language`, `duration`, `text`, `segments[]`, `words[]`. |\n| `srt` | `application/x-subrip` | SubRip subtitle file, one cue per VAD-segmented chunk. |\n| `vtt` | `text/vtt` | WebVTT subtitle file, one cue per VAD-segmented chunk. |\n\n`json` shape:\n```json\n{ \"text\": \" full transcript as a single string\" }\n```\n\n`verbose_json` shape — `segments` and `words` are always present (empty arrays for backends with no alignment output):\n```json\n{\n  \"task\": \"transcribe\",\n  \"language\": \"en\",\n  \"duration\": 6.42,\n  \"text\": \" full transcript\",\n  \"segments\": [{ \"id\": 0, \"start\": 0.0, \"end\": 2.31, \"text\": \" ...\", \"tokens\": [], \"temperature\": 0.0, \"avg_logprob\": null, \"compression_ratio\": null, \"no_speech_prob\": null }],\n  \"words\": [{ \"word\": \" the\", \"start\": 0.0, \"end\": 0.12 }]\n}\n```\n\nWhisper-only confidence fields (`avg_logprob`, `compression_ratio`, `no_speech_prob`) are emitted as `null` regardless of backend so clients reading them don't crash. `tokens` is always `[]`.\n\nThe `sherpa` and `vosk` executors add a per-word `confidence` in the 0–1 range to each entry in `words` — Vosk reports the decoder's own score, Sherpa derives one from the model's per-token acoustic log-probabilities. No other backend emits it, so treat the field as optional.\n\n### Stereo Diarization\n\nPass `diarization=true` and upload a 2-channel file. Left channel = speaker `L`, right channel = speaker `R`. Each channel is transcribed independently, the two timelines are merged chronologically by segment start time.\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@interview-stereo.wav\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"diarization=true\" \\\n  -F \"response_format=verbose_json\" | jq\n```\n\nWhat changes:\n- `verbose_json` — every segment/word gets `\"channel\": \"L\"` or `\"R\"`. Segments re-numbered after merge.\n- `text` / `response_format=text` — rebuilt as alternating turn lines: `L: ...\\nR: ...\\n...`. Consecutive same-channel segments collapsed into one line per turn.\n- `srt` / `vtt` — each cue prefixed with `L:` / `R:`.\n\nCaveats:\n- Exactly **2 channels** required. Mono → 400. >2 channels → 400.\n- Latency ~2× the mono case (model runs sequentially on each channel).\n- The technique is exact for true two-mic setups (interview rigs, podcast splits). It does NOT magically separate speakers from a single-mic recording that's been rendered to stereo.\n\n### Translation\n\nCanary multitask models can translate speech → text in a non-source language. `canary-1b-flash` covers en↔de, en↔fr, en↔es. **The task is baked into the model slug**, not passed per-request — you add a translation-specific slug via custom `models.json` (see [Customizing the model registry](references/setup.md#customizing-the-model-registry)):\n\n```json\n{\n  \"models\": {\n    \"canary-1b-flash-de2en\": {\n      \"repo\": \"nvidia/canary-1b-flash\",\n      \"executor\": \"canary_multitask\",\n      \"default_source_lang\": \"de\",\n      \"default_target_lang\": \"en\",\n      \"default_task\": \"s2t_translation\",\n      \"languages\": [\"de\"]\n    }\n  }\n}\n```\n\nThen call it normally — `text` carries the English translation:\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@german-clip.wav\" \\\n  -F \"model=canary-1b-flash-de2en\" | jq\n```\n\n`canary-180m-flash` is English-ASR-only — don't point a translation slug at it. `canary-qwen-2.5b` is English ASR only too.\n\n### Long Files + VAD Chunking\n\nAudio longer than 30 s (`TALKIES_VAD_CHUNK_THRESHOLD`) gets sliced through Silero VAD into ≤28 s speech regions before being handed to the backend. Timestamps are re-assembled by offsetting each chunk's segment/word timings — you get one continuous `segments` list spanning the whole file.\n\nNo client-side change. Long files just work. Verify by checking `duration` in `verbose_json`.\n\n### Error Contract\n\n| Status | Shape | When |\n|---|---|---|\n| 200 | per `response_format` | success |\n| 400 | `{\"detail\": \"...\"}` | bad audio, mono+diarization, >2 ch+diarization, both/neither of `file`/`file_path`, invalid file_path, URL download failure (DNS, HTTP error, size exceeded, SSRF blocked) |\n| 401 | `{\"detail\": \"...\"}` | only when `TALKIES_AUTH_TOKEN` is set: missing/wrong bearer. Includes `WWW-Authenticate: Bearer`. |\n| 404 | `{\"detail\": \"...\"}` | unknown model slug, `file_path` references missing file, model-evict on an unloaded model, file op on a missing `/v1/files` path |\n| 413 | `{\"detail\": \"...\"}` | upload exceeded `TALKIES_MAX_UPLOAD_BYTES` (multipart `file` and `PUT /v1/files/{path}` only — not `file_path` URL) |\n| 422 | `{\"detail\": [...]}` | Pydantic validation (missing fields, wrong types) |\n| 500 | `{\"detail\": \"...\"}` | unhandled backend failure |\n\n## API — `POST /v1/audio/speech` (TTS)\n\nJSON body (not multipart). Returns the encoded audio bytes in the body with the matching `Content-Type` — no JSON envelope.\n\n**Every call sends your `input` text (and, for voice-cloning slugs, the referenced voice sample) to whatever `$TALKIES_URL` points at — that data leaves your host.** Point `$TALKIES_URL` only at a talkies instance you run or explicitly trust; prefer HTTPS when it's not localhost/LAN. Don't synthesize sensitive or confidential text through a server you don't control.\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"kokoro-82m\",\n        \"input\": \"The quick brown fox jumps over the lazy dog.\",\n        \"voice\": \"af_heart\",\n        \"response_format\": \"mp3\",\n        \"speed\": 1.0\n      }' \\\n  --output fox.mp3\n```\n\n### Request Body\n\n| Field | Required | Default | Notes |\n|---|---|---|---|\n| `model` | yes | — | TTS model slug. Kokoro: `kokoro-82m`, `kokoro-82m-nvidia`. Qwen3-TTS: `qwen3-tts-0.6b`, `qwen3-tts-1.7b` (base/cloning), `qwen3-tts-0.6b-custom`, `qwen3-tts-1.7b-custom` (preset speakers), `qwen3-tts-1.7b-design` (voice from NL description). Chatterbox: `chatterbox-turbo` (English only). Unknown → 404. ASR slug → 400. |\n| `input` | yes | — | Text to synthesize. Empty / whitespace-only → 400. No fixed length cap; for very long inputs split client-side. |\n| `voice` | no | model `default_voice` | Semantics shift per Qwen3 mode (see [Qwen3-TTS Modes](#qwen3-tts-modes)). Kokoro: voice name (default `af_heart`). Qwen3 `base`: path of a reference WAV (default `alloy`). Qwen3 `custom_voice`: one of the 9 preset speakers (default `Vivian`). Qwen3 `voice_design`: ignored — sentinel `\"design\"`. `chatterbox-turbo`: `builtin` (default) or the name of a `.wav` under `/data/custom-voices/`, longer than 5 seconds — shorter clips → 400. Unknown → 400 with catalog listed. |\n| `response_format` | no | `mp3` | `mp3` / `opus` / `aac` / `flac` / `wav` / `pcm`. |\n| `speed` | no | `1.0` | Playback rate, Kokoro only. Clamped to `[0.25, 4.0]`. **Ignored** by every Qwen3-TTS slug (no speed control in Qwen3-TTS) and by `chatterbox-turbo`. |\n| `instructions` | no | — | Free-form style prompt. **Required** for `qwen3-tts-1.7b-design` (the NL voice description; empty → 400). **Honoured** by Qwen3-TTS `base` mode and `qwen3-tts-1.7b-custom` (threaded as `instruct`). **Dropped** by `qwen3-tts-0.6b-custom` (0.6B CustomVoice checkpoint limitation — logs a WARNING) and both Kokoro slugs (no instruction input). Accepted on every slug for OpenAI parity. |\n| `language` | no | model `default_language` (`English`) | **Non-OpenAI extra field** (send via `extra_body={\"language\": \"...\"}` on official SDKs). Selects the spoken language for Qwen3 `custom_voice` / `voice_design`; `base` mode reads it from the voice's sibling `.lang` file. Silently ignored by Kokoro. |\n| `temperature` | no | `0.9` | **Non-OpenAI extra, Qwen3-TTS only** (`extra_body`). Sampler temperature, `[0.0, 2.0]`. Ignored by Kokoro. |\n| `top_k` | no | `50` | **Non-OpenAI extra, Qwen3-TTS only.** Top-k truncation, `[1, 1000]`. Ignored by Kokoro. |\n| `top_p` | no | `1.0` | **Non-OpenAI extra, Qwen3-TTS only.** Nucleus sampling, `[0.0, 1.0]`. Ignored by Kokoro. |\n| `repetition_penalty` | no | `1.05` | **Non-OpenAI extra, Qwen3-TTS only.** Penalizes codec-token repeats, `[0.5, 2.0]`. Ignored by Kokoro. |\n| `max_new_tokens` | no | `2048` (model max) | **Non-OpenAI extra, Qwen3-TTS only.** Codec-step cap, `[1, 2048]`. Ignored by Kokoro. |\n| `do_sample` | no | `true` | **Non-OpenAI extra, Qwen3-TTS only.** `false` = greedy decode. Ignored by Kokoro. |\n\nOut-of-range sampling values → 422 (Pydantic validation).\n\n### Qwen3-TTS Modes\n\nThe Qwen3-TTS mode is implicit in the model slug — the OpenAI wire format stays pure (`model` / `voice` / `instructions` / `input`), with `voice` and `instructions` carrying mode-specific semantics. No new endpoints.\n\n| Mode | Slugs | What `voice` means | What `instructions` means |\n|---|---|---|---|\n| **base** (voice cloning) | `qwen3-tts-0.6b`, `qwen3-tts-1.7b` | Path of a reference `.wav` under the voices dirs (`.wav` stripped) | Optional style hint (passed as `instruct`) |\n| **custom_voice** (preset speakers) | `qwen3-tts-0.6b-custom`, `qwen3-tts-1.7b-custom` | One of 9 preset speaker names | Emotion / style cue — 1.7B honours it; 0.6B drops it (checkpoint limitation, logs a WARNING) |\n| **voice_design** (NL description) | `qwen3-tts-1.7b-design` | Ignored — sentinel `\"design\"` | **Required.** NL description of the voice (e.g. \"A warm, friendly young female voice\"). Empty → 400. |\n\nThe 9 `custom_voice` preset speakers (also returned by `GET /v1/audio/voices` for the chosen slug): `Vivian`, `Serena`, `Uncle_Fu`, `Dylan`, `Eric` (Chinese), `Ryan`, `Aiden` (English), `Ono_Anna` (Japanese), `Sohee` (Korean).\n\n### Output Formats\n\n`response_format` picks the encoder applied to Kokoro's raw 24 kHz mono PCM. ffmpeg does the conversion in-process; no temp files.\n\n| `response_format` | Content-Type | Codec / container | Notes |\n|---|---|---|---|\n| `mp3` (default) | `audio/mpeg` | libmp3lame, 128 kbps CBR | Most universal. |\n| `opus` | `audio/ogg` | libopus, 64 kbps VBR, Ogg container | Best quality-per-byte for speech. |\n| `aac` | `audio/aac` | AAC-LC, 128 kbps, ADTS | iOS-friendly. |\n| `flac` | `audio/flac` | FLAC | Lossless. |\n| `wav` | `audio/wav` | PCM s16le, 24 kHz mono, RIFF header | Lossless, largest. |\n| `pcm` | `application/octet-stream` | Raw PCM s16le, 24 kHz mono — no container, no header | Real-time chaining. Caller must know sample rate / format. |\n\n### Streaming PCM (Qwen3-TTS)\n\n`qwen3-tts-*` slugs stream `response_format=pcm` requests as chunked int16 PCM instead of buffering the whole utterance — first-audio latency drops from seconds to sub-second. Kokoro and every non-`pcm` format are unaffected (buffered as normal). Response carries an `X-Sample-Rate` header; chunk size (codec steps per yielded chunk) is tunable via `TALKIES_QWEN3_STREAM_CHUNK_SIZE` (default 8) — see [references/setup.md](references/setup.md).\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"qwen3-tts-0.6b\",\"input\":\"Streaming test.\",\"response_format\":\"pcm\"}' \\\n  --output stream.pcm\n```\n\n### Voices\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/voices | jq\n```\n\nReturns `{\"voices\": [{\"voice\", \"model\", \"default\", \"origin\"}]}`. The `origin` field is only present for engines that distinguish baked-in vs user-supplied voices (currently `qwen3-tts-0.6b` — `\"builtin\"` for image-baked samples, `\"custom\"` for `/data/custom-voices/` mounts). Kokoro entries omit `origin`.\n\n**Kokoro voices** encode `<lang_code><gender>_<name>`:\n\n| Prefix | Language |\n|---|---|\n| `af_` / `am_` | American English (female / male) |\n| `bf_` / `bm_` | British English (female / male) |\n| `ef_` / `em_` | Spanish |\n| `ff_` | French |\n| `hf_` / `hm_` | Hindi |\n| `if_` / `im_` | Italian |\n| `pf_` / `pm_` | Portuguese (Brazilian) |\n\n41 voices ship in the image. Japanese (`jf_*` / `jm_*`) and Chinese (`zf_*` / `zm_*`) are filtered out because they need the optional `misaki[ja]` / `misaki[zh]` extras (MeCab + pypinyin chains).\n\n**Qwen3-TTS voices** come from two on-disk dirs merged into one catalog:\n\n- `/opt/talkies/qwen3-voices/` — baked into the CUDA image. Ships three curated samples (`alloy`, `echo`, `fable`) so voice cloning works out-of-the-box. `origin=builtin`.\n- `/data/custom-voices/` — host-mounted via the data volume. Drop `foo/bar/me.wav` and voice `foo/bar/me` immediately appears in `GET /v1/audio/voices` (catalog is rescanned per request — no restart). `origin=custom`.\n\nVoice names are the wav's path relative to its parent dir with `.wav` stripped — nested subdirs are preserved. `custom-voices/team-a/jane.wav` → voice `team-a/jane`. Custom voices shadow builtin voices with the same name; dropping a `custom-voices/alloy.wav` overrides the builtin `alloy` sample (its `origin` flips to `custom`).\n\nOptional sibling metadata next to each `<name>.wav`:\n- `<name>.txt` — reference transcript for the clip (ICL voice cloning works without it, but clone fidelity is noticeably better with a faithful transcript).\n- `<name>.lang` — language label string (defaults to `English`).\n\nPath-traversal guard: hostile symlinks whose `resolve()` escapes the voices dir are skipped (the wav can't be used to read arbitrary host files as a voice prompt).\n\n```bash\n# Add a custom clone voice (server picks it up on next request — no restart).\nmkdir -p ~/talkies-data/custom-voices/team-a\ncp jane-reading.wav ~/talkies-data/custom-voices/team-a/jane.wav\necho \"And the silken sad uncertain rustling of each purple curtain.\" \\\n  > ~/talkies-data/custom-voices/team-a/jane.txt\n\n# Use it.\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"qwen3-tts-0.6b\",\n        \"input\": \"Hello from a cloned voice.\",\n        \"voice\": \"team-a/jane\",\n        \"response_format\": \"wav\"\n      }' \\\n  --output cloned.wav\n```\n\n**First synth is slow** on Qwen3-TTS — the predictor + talker CUDA graphs are captured on first call (~30-60 s on a mid-range GPU). Subsequent generations are sub-second. The model and graphs stay resident until evicted by sibling load or the idle sweeper.\n\n### Error Contract (TTS)\n\n| Status | When |\n|---|---|\n| 200 | success (audio bytes in body) |\n| 400 | empty `input`, unknown `voice`, unsupported `response_format`, model isn't TTS (e.g. POSTing `whisper-large-v3` here) |\n| 401 | `TALKIES_AUTH_TOKEN` set, missing / wrong bearer |\n| 404 | unknown `model` slug |\n| 422 | Pydantic validation (missing required fields, wrong types) |\n| 500 | unhandled ffmpeg or kokoro internal failure |\n| 503 | TTS snapshot files missing under `${TALKIES_DATA_DIR}/models/<slug>/` (slug excluded from `TALKIES_ENABLED_MODELS` but still being called); or `qwen3-tts-0.6b` requested on a non-CUDA device (the backend hard-fails at load time) |\n\n## Resource-Management Endpoints (Ollama-Style)\n\ntalkies mirrors a subset of [speaches](https://github.com/speaches-ai/speaches) / Ollama, so a LiteLLM proxy can drive both.\n\n| Endpoint | Behavior |\n|---|---|\n| `GET /healthz` | Unauthenticated liveness. Returns `{ok, device, models}`. |\n| `GET /v1/models` | OpenAI-style list of configured slugs. Each entry includes a `modality` field (`asr` or `tts`) so clients can filter. |\n| `GET /api/ps` | Currently-loaded models with per-model `idle_seconds`. |\n| `DELETE /api/ps/{model_id}` | Evict one model from memory. Slug can be URL-encoded (`/` → `%2F`). 404 if not loaded. |\n| `POST /unload` | Evict every loaded model. Returns the list actually unloaded. |\n\nModel eviction (`DELETE /api/ps/...`, `POST /unload`) forces a cold-load for anyone mid-request, so only do it for explicit maintenance the user asked for (e.g. \"free up VRAM\"). Auth is only enforced when `TALKIES_AUTH_TOKEN` is set — require it on shared deployments and don't expose these routes on an unauthenticated network.\n\nBehind these: an **idle sweeper** runs every `TALKIES_SWEEPER_INTERVAL` s (default 60) and unloads anything not used in `TALKIES_MODEL_TTL` s (default 600). Set `TALKIES_MODEL_TTL=0` to disable.\n\nThere's also **sibling eviction at request time** — every transcribe or speech request evicts other loaded models so VRAM doesn't get split. ASR and TTS share the same pool; loading Kokoro evicts a resident Whisper and vice versa. One model resident at a time, per container. If you need two models simultaneously, run two containers.\n\n```bash\n# Which models are loaded right now.\ncurl -s $TALKIES_URL/api/ps | jq\n\n# Free VRAM after a job — evict one model.\ncurl -s -X DELETE \"$TALKIES_URL/api/ps/whisper-large-v3-turbo\"\n\n# Or evict everything.\ncurl -s -X POST $TALKIES_URL/unload | jq\n```\n\n## Server-Side File Staging (`/v1/files`)\n\nFor repeated transcribes of the same file (different `response_format`, different model, iterating on params), stage the file once and reference it by path. Files land under `${TALKIES_DATA_DIR}/files/<path>`.\n\n**Staged files persist until explicitly removed** — nothing auto-expires them — and `GET /v1/files` enumerates every staged path to anyone who can reach the API. Don't stage sensitive/private media on a server without auth enabled (`TALKIES_AUTH_TOKEN`); clean up staged files when done.\n\n**Guardrail — this is a shared, unisolated bucket, not a private workspace.** There's no per-caller ownership: any path any caller staged is listable and readable by any other caller with API access. An agent must:\n- only read or delete paths it staged itself in the current, user-approved workflow;\n- never call `GET /v1/files` to browse/enumerate what other callers have staged, and never delete a path it didn't create;\n- clean up what it staged once the workflow is done, since nothing expires automatically.\n\nAt the deployment level: require `TALKIES_AUTH_TOKEN` by default, treat per-caller isolation and retention limits as the operator's responsibility (talkies itself provides neither).\n\n| Endpoint | Behavior |\n|---|---|\n| `GET /v1/files` | List every staged file. Returns `{\"files\": [{\"path\", \"size\", \"modified\"}]}`. **Enumerable by anyone with API access — no per-file ownership/isolation.** |\n| `PUT /v1/files/{path}` | Upload raw bytes (`--data-binary @local-file`). Capped at `TALKIES_MAX_UPLOAD_BYTES`. Atomic write (`.part` → rename). |\n| `GET /v1/files/{path}` | Streams file back. Content-Type guessed by extension. 404 if missing. |\n| `DELETE /v1/files/{path}` | Removes a staged file and prunes empty parent dirs (404 if missing). Files don't self-expire, so call this to clean up when done. On a shared bucket every path is listable by any caller, so only remove paths you staged yourself. |\n\n```bash\n# Stage once.\ncurl -X PUT --data-binary @lecture.mp3 \\\n  -H \"Content-Type: audio/mpeg\" \\\n  $TALKIES_URL/v1/files/lectures/2026-03-15/lecture.mp3\n\n# Reuse across multiple transcribe calls.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=lectures/2026-03-15/lecture.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"response_format=verbose_json\" | jq\n\n# Cleanup.\ncurl -X DELETE $TALKIES_URL/v1/files/lectures/2026-03-15/lecture.mp3\n```\n\nPath safety: null bytes, backslashes, `.` / `..` segments and double slashes are rejected (400). Symlinks pointing outside the root are refused. Leading `/` is stripped — `/foo/bar.mp3` and `foo/bar.mp3` resolve identically.\n\n### URL `file_path` (Download + Cache)\n\n`file_path` also accepts `http://` / `https://` URLs. First request downloads to `${TALKIES_DATA_DIR}/files/downloads/<sha256(url)[:16]>-<basename>`, subsequent requests with the same URL hit the cache.\n\n**The download happens server-side, and the result is cached persistently on the talkies server's disk** — not a transient client-side fetch. Anyone who can reach the API can later list/read that cached copy via `GET /v1/files` (see [Server-Side File Staging](#server-side-file-staging-v1files) above). Don't pass URLs to private/sensitive media unless the talkies server itself is trusted and access-controlled (`TALKIES_AUTH_TOKEN`); invalidate the cache entry when done.\n\n```bash\n# First call: downloads, transcribes off the cached copy.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq\n\n# Second call: same URL → cache hit, no re-download.\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=canary-1b-flash\" \\\n  -F \"response_format=srt\" > ep-042.srt\n```\n\nDownloads appear in `GET /v1/files` listings under `downloads/`. Invalidate a single cached URL by removing it from `/v1/files/downloads/`.\n\nConstraints applied during download:\n- Size capped by `TALKIES_MAX_DOWNLOAD_BYTES` (default 1 GiB).\n- 5 redirect hops max; SSRF guard re-applied at every hop.\n- 10 s connect, 300 s per-chunk read timeout.\n- SSRF off by default. Set `TALKIES_BLOCK_PRIVATE_DOWNLOADS=true` to reject URLs whose hostname resolves to private/loopback/link-local/multicast/reserved IPs.\n\n## MCP Endpoint (`/v1/mcp`)\n\ntalkies exposes a [Model Context Protocol](https://modelcontextprotocol.io) server over Streamable HTTP at `/v1/mcp`. Same FastAPI process, same `BACKENDS` / `REGISTRY`, same auth middleware — a model loaded by the MCP `transcribe` tool is the same instance the HTTP endpoint sees.\n\nMCP exposes the ASR surface only. TTS (`/v1/audio/speech`) is HTTP-only — generated audio bytes don't round-trip through JSON-RPC cleanly. `list_models` filters out TTS slugs so `transcribe` only ever sees ASR backends.\n\n| Tool | What it does |\n|---|---|\n| `list_models` | Discover ASR slugs (TTS slugs are filtered out). Returns `[{slug, executor, default_source_lang, default_target_lang, default_task, loaded}]`. |\n| `transcribe` | Run ASR on a `file_path` (URL or staged path). Args: `model`, `language?`, `response_format?` (`json`/`verbose_json`/`text`/`srt`/`vtt`), `diarization?`. JSON formats return a JSON-encoded string; text/srt/vtt return raw. |\n| `list_files` | Same payload as `GET /v1/files`. |\n| `put_file` | Upload to staging. Body is base64 (`content_base64`). Decoded size capped at `TALKIES_MAX_UPLOAD_BYTES`. **For big files, prefer `PUT /v1/files/{path}` over HTTP** — JSON-RPC + base64 chews token budget. |\n| `get_file` | Read a staged file as base64. Same size cap. Same advice — for big bytes, hit `GET /v1/files/{path}` over HTTP. |\n| `delete_file` | Remove a staged file, prune empty parents. |\n\nThe transport requires `Accept: application/json, text/event-stream`. Wire it into Claude Code:\n\n```bash\nclaude mcp add --transport http talkies $TALKIES_URL/v1/mcp\n```\n\nWith auth:\n\n```bash\nclaude mcp add --transport http talkies $TALKIES_URL/v1/mcp \\\n  --header \"Authorization: Bearer $TALKIES_AUTH_TOKEN\"\n```\n\nNote: the canonical mount path is `/v1/mcp/` (trailing slash). Bare `/v1/mcp` is rewritten internally to `/v1/mcp/` so clients that don't follow Starlette's 307 redirect work too.\n\n### Raw JSON-RPC\n\nFor debugging or non-MCP-aware callers, hit it as JSON-RPC over HTTP POST:\n\n```bash\n# tools/list\ncurl -s $TALKIES_URL/v1/mcp/ \\\n  -H \"Content-Type: application/json\" \\\n  -H \"Accept: application/json, text/event-stream\" \\\n  -d '{\"jsonrpc\": \"2.0\", \"id\": 1, \"method\": \"tools/list\"}'\n\n# tools/call\ncurl -s $TALKIES_URL/v1/mcp/ \\\n  -H \"Content-Type: application/json\" \\\n  -H \"Accept: application/json, text/event-stream\" \\\n  -d '{\n    \"jsonrpc\": \"2.0\", \"id\": 2, \"method\": \"tools/call\",\n    \"params\": {\n      \"name\": \"transcribe\",\n      \"arguments\": {\n        \"file_path\": \"https://example.com/clip.mp3\",\n        \"model\": \"whisper-large-v3-turbo\",\n        \"response_format\": \"json\"\n      }\n    }\n  }'\n```\n\n## Bearer-Token Auth\n\nIf `TALKIES_AUTH_TOKEN` is set on the server, every route except `/healthz` and CORS preflight (`OPTIONS`) requires `Authorization: Bearer <token>`. Wrong/missing token returns 401 with `WWW-Authenticate: Bearer`. Compared with `hmac.compare_digest` (constant-time).\n\n```bash\ncurl -H \"Authorization: Bearer $TALKIES_AUTH_TOKEN\" $TALKIES_URL/v1/models\n```\n\nEmpty / unset token = wide open. For untrusted networks, combine the token with a reverse proxy doing TLS + rate limiting.\n\n## Typical Workflows\n\n### Quick one-off transcribe\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq -r .text\n```\n\n### Generate subtitles for a video\n\n```bash\nffmpeg -i video.mp4 -vn -acodec libmp3lame audio.mp3\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3\" \\\n  -F \"response_format=srt\" > video.srt\n# burn in:  ffmpeg -i video.mp4 -vf subtitles=video.srt -c:a copy video-subbed.mp4\n```\n\n### Iterate on the same file with different settings\n\n```bash\n# Stage once.\ncurl -X PUT --data-binary @lecture.mp3 \\\n  -H \"Content-Type: audio/mpeg\" \\\n  $TALKIES_URL/v1/files/work/lecture.mp3\n\n# Try different models / formats without re-uploading.\nfor fmt in json verbose_json srt; do\n  curl -s $TALKIES_URL/v1/audio/transcriptions \\\n    -F \"file_path=work/lecture.mp3\" \\\n    -F \"model=whisper-large-v3-turbo\" \\\n    -F \"response_format=$fmt\" > \"lecture.$fmt\"\ndone\n\n# Clean up the staged file when done (see the /v1/files reference above).\n```\n\n### Diarized interview transcript\n\n```bash\ncurl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@interview-stereo.wav\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"diarization=true\" \\\n  -F \"response_format=text\"\n# stdout:\n#   L: hi how's it going\n#   R: not bad you\n#   L: cool man\n```\n\n### Synthesize speech from text\n\n```bash\n# Default voice, MP3 output.\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"kokoro-82m\",\"input\":\"Greetings, human.\"}' \\\n  --output greetings.mp3\n\n# Pick a voice from GET /v1/audio/voices, choose a format.\ncurl -s $TALKIES_URL/v1/audio/voices | jq -r '.voices[].voice'\ncurl -s $TALKIES_URL/v1/audio/speech \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"model\": \"kokoro-82m\",\n        \"input\": \"Buongiorno, mondo.\",\n        \"voice\": \"if_sara\",\n        \"response_format\": \"opus\"\n      }' \\\n  --output ciao.opus\n```\n\n### Free VRAM after a job\n\n```bash\ncurl -s -X POST $TALKIES_URL/unload | jq\n```\n\n### Bulk transcribe from URLs\n\n```bash\nfor url in $(cat urls.txt); do\n  curl -s $TALKIES_URL/v1/audio/transcriptions \\\n    -F \"file_path=$url\" \\\n    -F \"model=whisper-large-v3-turbo\" \\\n    -F \"response_format=text\"\n  echo \"---\"\ndone\n```\n\nThe first hit on each URL downloads + caches; re-running the loop is free.\nFor local files, replace the `file_path` form field with `file=@path/to/audio`\nin the same request shape.\n\n## Tips\n\n1. **Use `whisper-large-v3-turbo`** as your default — it's the speed/quality sweet spot for general-purpose ASR. Switch to `whisper-large-v3` only when you need the last few % of accuracy on hard audio.\n2. **URL `file_path` over multipart upload** — if the audio is already at a URL, send the URL. Saves bandwidth (the file isn't going up and then back down), gets cached server-side, no upload size cap.\n3. **Stage repeated files** via `PUT /v1/files/{path}` and call with `file_path=` to avoid re-uploading on every retry/iteration.\n4. **`response_format=text`** for the \"just give me the string\" case — no `jq -r .text` needed, content-type is `text/plain`.\n5. **One active model family at a time** — every transcribe request evicts other loaded models. Multiple live streams may share one pinned model up to `TALKIES_STREAM_MAX_CONNECTIONS`; a request or stream for a different model receives a conflict until the streams end. Use two containers if you need concurrent models.\n6. **`POST /unload` after a job** — explicit eviction frees VRAM/RAM faster than waiting for the 10-min idle sweeper. Useful in CI / batch scripts.\n7. **`canary-qwen-2.5b` has no timestamps** — `verbose_json.segments` / `.words` come back empty, `srt`/`vtt` collapse to one cue. Use a Whisper or Canary multitask slug if you need timing data.\n8. **Diarization requires true stereo** — if your \"stereo\" file is the same mono signal copied to both channels, diarization won't separate speakers. The technique is exact for two-mic setups, useless otherwise.\n9. **Long files just work** — VAD chunking happens transparently. Don't pre-split. Send the whole file.\n10. **ASR's `prompt` / `temperature` are ignored** even though the request schema accepts them. TTS's `instructions` is different — Kokoro ignores it, but Qwen3-TTS honors it in `base` mode and `*-custom` (except the 0.6B checkpoint) and requires it for `*-design`.\n11. **Watch `/api/ps`** to see what's resident. A request that hangs at \"loading model\" is doing the first cold load — subsequent calls are fast.\n12. **Customizing the model registry** for translation slugs or to restrict the served set — see [references/setup.md](references/setup.md#customizing-the-model-registry).\n13. **Kokoro uses native voice names** — no OpenAI aliases. Hit `GET /v1/audio/voices` once to discover what's shipped; pass the `voice` field accordingly. The 41 voices cover en (US + UK), es, fr, hi, it, pt; ja/zh are filtered out.\n14. **Voice cloning is `qwen3-tts-0.6b`** — drop a `.wav` (10-30 s of clean speech is plenty) into `/data/custom-voices/<anywhere>.wav`. Optionally drop a sibling `.txt` with a faithful transcript for higher clone fidelity. The voice appears in `GET /v1/audio/voices` on the next request — no restart. CUDA required.\n15. **Qwen3-TTS first synth is slow** — CUDA graph capture runs once after model load (~30-60 s). Subsequent synths are sub-second. If you're benchmarking, throw away the first call.\n16. **Qwen3-TTS ignores `speed`** — the model has no playback-rate control. Pass it for OpenAI compat; nothing happens. Only Kokoro honors `speed`.\n17. **Both TTS engines emit 24 kHz mono PCM** — Kokoro and Qwen3-TTS both output int16 24 kHz mono. ffmpeg re-encodes into your chosen `response_format`; `pcm` (raw, no container) hands back that same rate directly — check the `X-Sample-Rate` response header on Qwen3-TTS streaming responses if you need it confirmed per-request.\n18. **TTS `response_format=pcm` is for chaining** — raw int16 mono PCM, no container, no header. Use it when piping into another encoder or a real-time playback path. Otherwise stick with `mp3` (default) or `opus` for size.\n19. **TTS evicts loaded ASR and vice versa** — they share the same one-model-resident pool. Synthesizing with Kokoro after a transcribe burst incurs Kokoro's cold load. Same applies to Qwen3-TTS (plus the CUDA-graph capture re-runs on cold reload).\n\nFile v1.3.17:_meta.json\n\n{\n  \"ownerId\": \"kn79dhvmpjng4rp2jjk8k0v5xx80ccbk\",\n  \"slug\": \"talkies\",\n  \"version\": \"1.3.17\",\n  \"publishedAt\": 1787306554940\n}\n\nFile v1.3.17:references/setup.md\n\n# talkies setup\n\n## Requirements\n\n- Docker\n- `linux/amd64` host (no arm64 images — `nemo_toolkit[asr]` + chain doesn't resolve cleanly on aarch64)\n- Optional: NVIDIA GPU + NVIDIA Container Toolkit for the CUDA image (required for `qwen3-tts-0.6b` voice cloning)\n- ~3 GB disk for the CPU image, ~11 GB for the CUDA image\n- Additional disk for selected model weights; set `TALKIES_ENABLED_MODELS` to avoid downloading the full registry\n- ~4 GB RAM minimum (whisper-large-v3 needs the working set + overhead); 12 GB+ VRAM for the GPU-only models\n\n## Quick Install\n\n### CPU\n\nServes 2× Whisper + `canary-180m-flash` + `nemotron-3.5-asr-0.6b` (CPU-optimized, via parakeet.cpp), four selectable Sherpa Zipformer variants, and `vosk-small-en-us-0.15` for ASR, plus `kokoro-82m` and `kokoro-82m-nvidia` for TTS. The CUDA-only ASR models aren't worth running on CPU, and the Qwen3-TTS family is CUDA-only.\n\n```bash\ndocker run -d --name talkies \\\n  -v $HOME/talkies-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest\n```\n\n### CUDA\n\nServes all twelve ASR models plus all three TTS engines / 4 backends (`kokoro-82m`, `kokoro-82m-nvidia`, the 5 Qwen3-TTS slugs, and `chatterbox-turbo`). Requires the NVIDIA Container Toolkit on the host.\n\nNemotron runs through the SHA-256-pinned upstream parakeet.cpp v0.5.0 CUDA 12\nbundle in this image, so both file transcription and native WebSocket sessions\nuse GPU offload. The matching CUDA 12.9 runtime libraries stay isolated under\n`/opt/parakeet` from the image's Python ML stack.\n\n```bash\ndocker run -d --name talkies \\\n  --gpus all \\\n  -v $HOME/talkies-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest-cuda\n```\n\nThe CUDA image expects `--gpus all`. Without a GPU assignment it retains its\n`TALKIES_DEVICE=cuda` image default, so model loading fails rather than silently\nfalling back. To use its CPU-compatible subset for debugging, explicitly set\n`-e TALKIES_DEVICE=cpu` and restrict `TALKIES_ENABLED_MODELS` to CPU-compatible\nslugs; GPU-only Qwen3-TTS slugs remain unavailable.\n\n**Verify:** `curl http://localhost:8000/healthz` returns `{\"ok\": true, \"device\": \"...\", \"models\": [...]}` once boot's done.\n\n**First boot:** the entrypoint downloads every enabled model into `/data/models/<slug>/` and creates `/data/files/` + `/data/custom-voices/`. Bind-mount `/data` so subsequent restarts are no-ops. Restrict the download set with `TALKIES_ENABLED_MODELS` to avoid pulling everything.\n\n## CPU vs CUDA Images\n\n| Image | Tag | Platforms | Models served | Image size |\n|---|---|---|---|---|\n| CPU | `psyb0t/talkies:latest` | `linux/amd64` | 2× Whisper, Canary-180m-Flash, Nemotron-3.5-ASR, Sherpa Zipformer ×4, Vosk, Kokoro-82M ×2 runtimes | ~3 GB |\n| CUDA | `psyb0t/talkies:latest-cuda` | `linux/amd64` | all twelve ASR + Kokoro-82M ×2 runtimes + Qwen3-TTS ×5 + Chatterbox Turbo | ~11 GB |\n\nThe CPU image only ships ASR models that actually finish in a sane time without a GPU. Parakeet-TDT is autoregressive (slow on CPU). Canary-1B and Canary-Qwen-2.5B need the CUDA image with `--gpus all`; use the CPU image for CPU workloads. Kokoro-82M ships in both images — at 82M params it synthesizes faster than real-time on a 4-core CPU, no GPU needed. Chatterbox Turbo is CUDA-only for the same reason as the heavier ASR models: it runs on CPU but measures roughly 5-10x slower than real-time, so it is not registered as a CPU slug.\n\nBoth images bake `espeak-ng` into the runtime layer because Kokoro's G2P for es/fr/hi/it/pt routes through it via `misaki.espeak.EspeakG2P`. The Python `kokoro==0.9.4` package and its lightweight dependency chain (`misaki`, no `[ja]` / `[zh]` extras) are pinned alongside the rest of the ML stack in `Dockerfile` / `Dockerfile.cuda`.\n\nThe CUDA image additionally bakes the `faster-qwen3-tts==0.2.6` MIT wrapper and three builtin Qwen3 reference voices (`alloy`, `echo`, `fable`) under `/opt/talkies/qwen3-voices/`. The model weights (`Qwen/Qwen3-TTS-12Hz-0.6B-Base`, Apache-2.0) are downloaded into `/data/models/qwen3-tts-0.6b/` at first boot like every other model.\n\nThe CUDA image also bakes `chatterbox-tts==0.1.7` (MIT) and `s3tokenizer==0.3.0` (Apache-2.0) from a separate hash-pinned `requirements-chatterbox.txt`, installed `--no-deps` because their declared dependency metadata conflicts with the image's pinned torch/transformers and pulls tooling that has no place in a runtime image. The `ResembleAI/chatterbox-turbo` weights (MIT, ungated) land in `/data/models/chatterbox-turbo/` at first boot. Its voices come from `/data/custom-voices/` plus a `builtin` speaker shipped inside the checkpoint — the Qwen3 reference voices are deliberately not shared with it, since several are shorter than its 5-second reference-clip minimum.\n\n## Environment Variables\n\n### Auth + bind\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_AUTH_TOKEN` | (empty = no auth) | Bearer token required on every route except `/healthz`. Empty/unset = wide open (historical default — fine on private networks). When set, `Authorization: Bearer <token>` required on every HTTP request AND every MCP call. Compared with `hmac.compare_digest`. |\n\nContainer binds `0.0.0.0:8000` unconditionally. Control network exposure at `docker run` time:\n- `-p 127.0.0.1:8000:8000` — loopback-only on the host.\n- `-p 8000:8000` — all host interfaces.\n- For untrusted networks, combine the token with a reverse proxy doing TLS + rate limiting.\n\n### Device + model registry\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_DEVICE` | image default (`cpu` CPU / `cuda` CUDA) | `auto` picks `cuda` if available else `cpu`; it is an accepted override. Pin to a specific GPU with `cuda:N`. |\n| `TALKIES_MODELS_FILE` | `/app/models.json` | Path to the model registry JSON. Override to ship a custom subset. The CPU image copies `models-cpu.json` to this path; the CUDA image copies `models.json` here. |\n| `TALKIES_ENABLED_MODELS` | (empty = all from `models.json`) | Comma-separated slug whitelist. Restricts both the boot-time snapshot download and the queryable surface of `/v1/models`. Unknown slugs fail fast on startup. |\n| `TALKIES_PRELOAD` | (empty) | Comma-separated slugs to load into RAM/VRAM at boot, before uvicorn accepts requests. Skips cold-load on first transcription. Must be a subset of `TALKIES_ENABLED_MODELS`. |\n| `TALKIES_MODEL_MAX_CONCURRENCY` | `1` | Fallback number of simultaneous inference requests admitted per model across HTTP, MCP, WebSocket ASR, buffered TTS, and streaming TTS. Registry `max_concurrency` values take precedence. |\n| `TALKIES_MODEL_CONCURRENCY` | (empty) | Comma-separated `model-slug=limit` overrides, for example `nemotron-3.5-asr-0.6b=2,kokoro-82m=4`. Unknown, disabled, duplicate, malformed, or out-of-range entries fail at startup. |\n\nEach registry model may define `max_concurrency` from 1 through 1024. The\nbundled Nemotron entry defaults to two in both images. Only one model may own\nactive inference slots at a time, which prevents sibling model eviction while\na request is still using its backend.\n\n### Data dir\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_DATA_DIR` | `/data` | Base data dir. Model snapshots → `$TALKIES_DATA_DIR/models/<slug>/` (flat per-model dirs, no HF cache layout). Staged uploads + URL downloads → `$TALKIES_DATA_DIR/files/`. Qwen3-TTS custom clone voices → `$TALKIES_DATA_DIR/custom-voices/` (nested subdirs preserved as voice names). Bind-mount to persist across restarts. |\n\n**Security note on `$TALKIES_DATA_DIR/files/`:** staged uploads and cached URL downloads persist here **indefinitely** — nothing auto-expires them — and are enumerable by any caller via `GET /v1/files` (no per-caller isolation; see [Server-Side File Staging](../SKILL.md#server-side-file-staging-v1files) in SKILL.md). This is a shared bucket: an agent must only read/delete paths it staged itself, must never enumerate or delete other callers' files, and should clean up after its own workflow. Deploy with `TALKIES_AUTH_TOKEN` set by default, add per-caller isolation and retention limits at the deployment/proxy level if the deployment isn't fully trusted, and least-privilege network exposure otherwise.\n\n### Lifecycle (idle sweeper + load timeouts)\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_MODEL_TTL` | `600` (10 min) | Idle time before a loaded backend is unloaded by the sweeper. Bare number = seconds; also accepts Go-style `3h30m5s`, `45m`, `90s`. `0` disables auto-unload. |\n| `TALKIES_SWEEPER_INTERVAL` | `60` | How often the sweeper checks for idle models. |\n| `TALKIES_LOAD_TIMEOUT` | `300` | Parsed configuration reserved for a future model-load timeout; the current server does not apply it. |\n\n### Upload + download caps\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_MAX_UPLOAD_BYTES` | `104857600` (100 MB) | Reject `POST /v1/audio/transcriptions` multipart `file` and `PUT /v1/files/{path}` bodies larger than this with 413. |\n| `TALKIES_MAX_DOWNLOAD_BYTES` | `1073741824` (1 GiB) | Abort URL downloads (when `file_path` is an http(s) URL) larger than this. Larger default because downloads stream straight to disk, no in-memory buffering. |\n| `TALKIES_BLOCK_PRIVATE_DOWNLOADS` | `false` | Set to `true` to refuse URL downloads whose hostname resolves to private/loopback/link-local/multicast/reserved IPs. Default `false` because the typical self-hosted deployment is a LAN box fetching from another LAN box. Flip to `true` if exposed to untrusted clients. |\n\n### VAD knobs\n\nAudio longer than `TALKIES_VAD_CHUNK_THRESHOLD` seconds gets sliced through Silero VAD into ≤`TALKIES_VAD_MAX_SPEECH`-second speech regions before being handed to the backend.\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_VAD_CHUNK_THRESHOLD` | `30.0` | Audio longer than this (seconds) goes through VAD chunking. Shorter clips skip it. |\n| `TALKIES_VAD_MAX_SPEECH` | `28.0` | Max length of a single VAD-detected speech region (seconds). Should stay under Whisper's 30 s internal window. |\n| `TALKIES_VAD_MIN_SILENCE_MS` | `500` | Silero VAD param — minimum gap (ms) to consider a region break. |\n| `TALKIES_VAD_SPEECH_PAD_MS` | `200` | Silero VAD param — silence padding (ms) around each detected speech region. |\n| `TALKIES_VAD_THRESHOLD` | `0.5` | Silero VAD speech-probability threshold. Lower = more aggressive. |\n\n### Live ASR streaming\n\n`WS /v1/audio/transcriptions/stream` accepts headerless 16 kHz mono PCM16LE.\nIt is separate from the OpenAI-compatible upload route. See the repository's\n[`docs/streaming.md`](https://github.com/psyb0t/docker-talkies/blob/main/docs/streaming.md)\nfor the protocol and client examples.\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_STREAM_MAX_CONNECTIONS` | `4` | Maximum active ASR WebSockets per container. Streams may share one pinned model; attempts to switch models while one is active return a conflict. |\n| `TALKIES_STREAM_MAX_FRAME_BYTES` | `65536` | Maximum binary PCM frame size. Frames must be non-empty, contain whole 16-bit samples, and be 2–16777216 bytes. |\n| `TALKIES_STREAM_MAX_BUFFER_SECONDS` | `5` | Faster-whisper rolling-window budget. Must hold one configured maximum-size frame; native decoders process each frame directly. |\n| `TALKIES_STREAM_IDLE_TIMEOUT` | `30s` | Maximum wait between client messages before close code 4408. |\n| `TALKIES_STREAM_MAX_DURATION` | `4h` | Maximum accepted audio duration per WebSocket. |\n\n### Qwen3-TTS streaming\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_QWEN3_STREAM_CHUNK_SIZE` | `8` | Codec steps decoded per yielded chunk when `response_format=pcm` streams from a `qwen3_tts` backend (~1 s of audio per 12 steps). Only relevant to that streaming path. |\n\n### Logging\n\n| Var | Default | What it does |\n|---|---|---|\n| `TALKIES_LOG_LEVEL` (falls back to `LOG_LEVEL`) | `info` | `debug` / `info` / `warn` / `error` / `fatal` (case-insensitive; `warning` / `critical` also accepted). Unrecognized values fail fast at startup. JSON structured logs on stdout. **`debug` logs full request/response bodies** (TTS input text, cloned-voice reference transcripts, ASR transcripts) — PII; a one-time WARNING fires at startup when active. |\n\n### Internal\n\n| Var | Default | What it does |\n|---|---|---|\n| `HF_HUB_OFFLINE` | `1` (in image) | Refuse network calls from HuggingFace Hub at runtime. The entrypoint transparently unsets it for the one-shot prefetch step so the initial download works; the server process itself runs offline. Don't touch unless debugging. |\n\n## Common Configurations\n\n```bash\n# Restrict to just the small/fast models (saves first-boot download time).\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_ENABLED_MODELS=whisper-large-v3-turbo,canary-180m-flash \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Preload at boot so the first request doesn't pay the cold-load tax.\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_ENABLED_MODELS=whisper-large-v3-turbo \\\n  -e TALKIES_PRELOAD=whisper-large-v3-turbo \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Bearer auth on a public-facing deployment.\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_AUTH_TOKEN=$(openssl rand -hex 32) \\\n  -e TALKIES_BLOCK_PRIVATE_DOWNLOADS=true \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Loopback only (rely on reverse proxy for external access).\ndocker run -d -p 127.0.0.1:8000:8000 \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Disable auto-unload (keep model resident forever).\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_MODEL_TTL=0 \\\n  -v $HOME/talkies-data:/data \\\n  psyb0t/talkies:latest\n\n# Bump upload + download caps for huge files.\ndocker run -d -p 8000:8000 \\\n  -e TALKIES_MAX_UPLOAD_BYTES=10737418\n\nArchive v1.3.16: 5 files, 31285 bytes\n\nFiles: references/setup.md (24535b), scripts/bulk_transcribe.sh (3497b), skill-card.md (3323b), SKILL.md (48555b), _meta.json (127b)\n\nArchive v1.3.15: 5 files, 30806 bytes\n\nFiles: references/setup.md (24510b), scripts/bulk_transcribe.sh (3497b), skill-card.md (2659b), SKILL.md (47716b), _meta.json (127b)\n\nArchive v1.3.14: 5 files, 30692 bytes\n\nFiles: references/setup.md (24291b), scripts/bulk_transcribe.sh (3497b), skill-card.md (2760b), SKILL.md (47716b), _meta.json (127b)\n\nArchive v1.3.13: 5 files, 29725 bytes\n\nFiles: references/setup.md (23415b), scripts/bulk_transcribe.sh (3497b), skill-card.md (2696b), SKILL.md (46226b), _meta.json (127b)\n\nArchive v1.3.12: 5 files, 29823 bytes\n\nFiles: references/setup.md (23129b), scripts/bulk_transcribe.sh (3738b), skill-card.md (3060b), SKILL.md (46301b), _meta.json (127b)\n\nArchive v1.3.11: 5 files, 29374 bytes\n\nFiles: references/setup.md (22405b), scripts/bulk_transcribe.sh (3738b), skill-card.md (2778b), SKILL.md (46301b), _meta.json (127b)\n\nArchive v1.3.10: 5 files, 29453 bytes\n\nFiles: references/setup.md (22405b), scripts/bulk_transcribe.sh (3738b), skill-card.md (3128b), SKILL.md (46019b), _meta.json (127b)\n\nArchive v1.3.9: 5 files, 29660 bytes\n\nFiles: references/setup.md (22341b), scripts/bulk_transcribe.sh (3738b), skill-card.md (3589b), SKILL.md (46283b), _meta.json (126b)","readmeExcerpt":"Skill: talkies Owner: psyb0t Summary: Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"export TALKIES_URL=http://localhost:8000"},{"language":"bash","snippet":"export TALKIES_AUTH_TOKEN=<your-token>\n# every request below needs: -H \"Authorization: Bearer $TALKIES_AUTH_TOKEN\""},{"language":"bash","snippet":"curl -s $TALKIES_URL/v1/models | jq"},{"language":"bash","snippet":"curl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq"},{"language":"bash","snippet":"curl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file_path=https://example.com/podcasts/ep-042.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" | jq"},{"language":"bash","snippet":"curl -s $TALKIES_URL/v1/audio/transcriptions \\\n  -F \"file=@audio.mp3\" \\\n  -F \"model=whisper-large-v3-turbo\" \\\n  -F \"response_format=verbose_json\" | jq"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: talkies\ndescription: Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 14 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk, plus wav2vec2 and ZIPA phoneme recognizers that emit IPA); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.\nhomepage: https://github.com/psyb0t/docker-talkies\nuser-invocable: true\npermissions:\n  network: \"Outbound HTTP to the configured TALKIES_URL; the Talkies server also fetches URLs supplied as file_path.\"\n  shell: \"Documented setup and workflow examples invoke local curl, ffmpeg, and docker commands.\"\n  filesystem: \"Reads and writes server-side staged files through /v1/files; the skill itself does not access the local filesystem.\"\nmetadata:\n  { \"openclaw\": { \"emoji\": \"🎙️\", \"primaryEnv\": \"TALKIES_URL\", \"requires\": { \"bins\": [\"docker\", \"curl\"] } } }\n---\n\n# talkies\n\nSelf-hosted speech service — ASR and TTS, one container. OpenAI-compatible wire shape on both endpoints; point an OpenAI client at it, change the model slug, done.\n\nASR (`POST /v1/audio/transcriptions`): fourteen bundled slugs — `whisper-large-v3`, `whisper-large-v3-turbo`, `parakeet-tdt-0.6b-v3`, `nemotron-3.5-asr-0.6b`, `canary-180m-flash`, `canary-1b-flash`, `canary-qwen-2.5b`, four selectable English Sherpa Zipformer variants, `vosk-small-en-us-0.15`, and two phoneme recognizers, `wav2vec2-xlsr-53-espeak` and `zipa-ipa`, that return the IPA phones spoken instead of words.\n\nTTS (`POST /v1/audio/speech`): 3 engines / 4 backends across 8 slugs — `kokoro-82m` (PyTorch) and `kokoro-82m-nvidia` (ONNX/ORT) with 41 baked voices across en/es/fr/hi/it/pt, plus the CUDA-only Qwen3-TTS family: `qwen3-tts-0.6b` / `qwen3-tts-1.7b` (voice cloning from reference clips), `qwen3-tts-0.6b-custom` / `qwen3-tts-1.7b-custom` (9 preset speakers), `qwen3-tts-1.7b-design` (voice from an NL description), plus the CUDA-only `chatterbox-turbo` (English only; 19 inline emotion tags; clones from a reference `.wav` with no transcript). Discover voices via `GET /v1/audio/voices`.\n\nExtras: live PCM ASR over WebSocket, stereo diarization on transcription, URL `file_path` fetching, server-side file staging, MCP endpoint with 6 ASR-side tools, optional bearer-token auth.\n\nFor installation, configuration, and container setup, see [references/setup.md](references/setup.md).\n\n## Security & safety\n\nThis skill is **not** low-risk to orchestrate blindly — it issues local shell commands (`curl`, `ffmpeg`, `docker` in the typical deployment/setup path) and outbound HTTP requests to whatever `$TALKIES_URL` points at:\n\n- **Outbound HTTP to an operator-chosen host —"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn79dhvmpjng4rp2jjk8k0v5xx80ccbk\",\n  \"slug\": \"talkies\",\n  \"version\": \"1.3.18\",\n  \"publishedAt\": 1787871031680\n}"},{"path":"references/setup.md","content":"# talkies setup\n\n## Requirements\n\n- Docker\n- `linux/amd64` host (no arm64 images — `nemo_toolkit[asr]` + chain doesn't resolve cleanly on aarch64)\n- Optional: NVIDIA GPU + NVIDIA Container Toolkit for the CUDA image (required for `qwen3-tts-0.6b` voice cloning)\n- ~3 GB disk for the CPU image, ~11 GB for the CUDA image\n- Additional disk for selected model weights; set `TALKIES_ENABLED_MODELS` to avoid downloading the full registry\n- ~4 GB RAM minimum (whisper-large-v3 needs the working set + overhead); 12 GB+ VRAM for the GPU-only models\n\n## Quick Install\n\n### CPU\n\nServes 2× Whisper + `canary-180m-flash` + `nemotron-3.5-asr-0.6b` (CPU-optimized, via parakeet.cpp), four selectable Sherpa Zipformer variants, and `vosk-small-en-us-0.15` for ASR, plus `kokoro-82m` and `kokoro-82m-nvidia` for TTS. The CUDA-only ASR models aren't worth running on CPU, and the Qwen3-TTS family is CUDA-only.\n\n```bash\ndocker run -d --name talkies \\\n  -v $HOME/talkies-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest\n```\n\n### CUDA\n\nServes all fourteen ASR models plus all three TTS engines / 4 backends (`kokoro-82m`, `kokoro-82m-nvidia`, the 5 Qwen3-TTS slugs, and `chatterbox-turbo`). Requires the NVIDIA Container Toolkit on the host.\n\nNemotron runs through the SHA-256-pinned upstream parakeet.cpp v0.5.0 CUDA 12\nbundle in this image, so both file transcription and native WebSocket sessions\nuse GPU offload. The matching CUDA 12.9 runtime libraries stay isolated under\n`/opt/parakeet` from the image's Python ML stack.\n\n```bash\ndocker run -d --name talkies \\\n  --gpus all \\\n  -v $HOME/talkies-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/talkies:latest-cuda\n```\n\nThe CUDA image expects `--gpus all`. Without a GPU assignment it retains its\n`TALKIES_DEVICE=cuda` image default, so model loading fails rather than silently\nfalling back. To use its CPU-compatible subset for debugging, explicitly set\n`-e TALKIES_DEVICE=cpu` and restrict `TALKIES_ENABLED_MODELS` to CPU-compatible\nslugs; GPU-only Qwen3-TTS slugs remain unavailable.\n\n**Verify:** `curl http://localhost:8000/healthz` returns `{\"ok\": true, \"device\": \"...\", \"models\": [...]}` once boot's done.\n\n**First boot:** the entrypoint downloads every enabled model into `/data/models/<slug>/` and creates `/data/files/` + `/data/custom-voices/`. Bind-mount `/data` so subsequent restarts are no-ops. Restrict the download set with `TALKIES_ENABLED_MODELS` to avoid pulling everything.\n\n## CPU vs CUDA Images\n\n| Image | Tag | Platforms | Models served | Image size |\n|---|---|---|---|---|\n| CPU | `psyb0t/talkies:latest` | `linux/amd64` | 2× Whisper, Canary-180m-Flash, Nemotron-3.5-ASR, Sherpa Zipformer ×4, Vosk, wav2vec2 + ZIPA phoneme, Kokoro-82M ×2 runtimes | ~3 GB |\n| CUDA | `psyb0t/talkies:latest-cuda` | `linux/amd64` | all fourteen ASR + Kokoro-82M ×2 runtimes + Qwen3-TTS ×5 + Chatterbox Turbo | ~11 GB |\n\nThe CPU image only ships ASR models that actually finish in a sane time without a GPU. Parakeet-TDT is autoregressive (slow on CPU). Canary-1"},{"path":"skill-card.md","content":"## Description:\n\ntalkies helps agents use a self-hosted OpenAI-compatible speech service for transcription, live ASR, speech synthesis, subtitles, server-side file staging, and MCP-based ASR workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[psyb0t](https://clawhub.ai/user/psyb0t)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agents use talkies to run speech-to-text, text-to-speech, subtitle generation, live transcription, and batch transcription workflows against a trusted self-hosted Talkies server. It is useful when an OpenAI-compatible audio API, optional MCP ASR tools, and local control over speech models are required.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Speech input, synthesized text, and voice-cloning reference samples are sent to the configured Talkies server.\n\nMitigation: Point TALKIES_URL only at an instance the operator runs or explicitly trusts, prefer HTTPS outside localhost or a protected LAN, and avoid sending sensitive media or text to untrusted servers.\n\nRisk: An unconfigured Talkies deployment can expose the API, staged files, and management endpoints to anyone who can reach the port.\n\nMitigation: Bind the service to localhost or a protected network, set TALKIES_AUTH_TOKEN for shared or public deployments, and use a reverse proxy with TLS and rate limiting when exposing it externally.\n\nRisk: Server-side URL file_path fetching can access operator-side network resources and persist downloaded media on the server.\n\nMitigation: Enable TALKIES_BLOCK_PRIVATE_DOWNLOADS before accepting untrusted callers, pass only trusted URLs, and delete cached downloads after the workflow is complete.\n\nRisk: Server-side staged files are persistent, shared, and enumerable by callers with API access.\n\nMitigation: Do not stage sensitive media on shared instances without authentication, only read or delete files created for the current workflow, and clean up staged files promptly.\n\nRisk: Voice cloning and expressive TTS can be misused to impersonate real people.\n\nMitigation: Clone or synthesize a real person's voice only with explicit authorization and consent, and remove custom voice samples when no longer needed.\n\nRisk: Setup and workflow examples execute local docker, curl, and ffmpeg commands, and latest container tags may change over time.\n\nMitigation: Review commands before execution, run them in an appropriate environment, and prefer pinned image digests for repeatable deployments.\n\nRisk: Debug logging can expose request and response content, including transcripts, TTS text, and voice reference transcripts.\n\nMitigation: Keep production deployments at info or higher log levels and use debug logging only for local troubleshooting with synthetic or disposable data.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/psyb0t/skills/talkies)\n- [talkies setup guide](references/setup.md)\n- [Talkies repository](https://github.com/psyb0t/docker-talkies"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1994,"uniquenessScore":40,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-09T22:59:11.240Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-09T22:59:11.240Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:40:04.942Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}