Audio Transcribe
This skill should be used when the user explicitly asks to "transcribe a meeting", "transcribe audio", "transcribe a meeting recording", "convert audio to te...
Rank
62
Safety
84
Downloads
1.2k
Updated
Oct 11, 2026
Version
1.7.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.2K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.7.1release · observed May 1, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr- Install using `clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/snapshot"
Documentation
CLAWHUB
149,258 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: audio-transcribe
version: 1.7.1
description: >
This skill should be used when the user explicitly asks to "transcribe a meeting",
"transcribe audio", "transcribe a meeting recording",
"convert audio to text", "generate meeting minutes from audio",
"do speech-to-text", "transcribe with speaker diarization",
"identify speakers in audio", "transcribe Chinese audio",
"transcribe English audio", "transcribe Japanese audio",
"multi-speaker transcription", "transcribe a podcast",
"transcribe podcast episode", "transcribe an interview",
"convert podcast to text", "podcast to transcript",
or mentions FunASR, Paraformer, SenseVoice, Whisper, MiMo, MiMo-V2.5-ASR,
meeting transcription, podcast transcription, or speaker diarization.
Supports multi-speaker meeting and podcast transcription in Chinese,
English, Japanese, Korean, Cantonese, and 99 languages (via Whisper),
plus Xiaomi MiMo-V2.5-ASR (8B, local GPU) for stronger proper-noun and
code-switching accuracy. Automatic speaker diarization via CAM++,
hotword biasing (FunASR path), LLM cleanup. FunASR works on GPU and CPU;
MiMo requires a local CUDA GPU with >=20GB VRAM.
metadata:
openclaw:
requires:
bins: ["python3", "ffmpeg"]
env_vars:
- name: AWS_REGION
required: false
description: "AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed."
- name: ANTHROPIC_API_KEY
required: false
description: "API key for Anthropic Claude LLM cleanup"
- name: OPENAI_API_KEY
required: false
description: "API key for OpenAI-compatible LLM cleanup"
- name: OPENAI_BASE_URL
required: false
description: "Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)"
emoji: "🎙️"
homepage: "https://github.com/zxkane/audio-transcriber"
---
# Meeting & Podcast Transcription (FunASR + MiMo)
Transcribe multi-speaker audio into structured Markdown with automatic
speaker diarization, hotword biasing, and optional LLM cleanup. Two
ASR engine families are available: **FunASR** (Paraformer / SenseVoice /
Whisper — fast, cheap, GPU or CPU, 99 languages) and **MiMo-V2.5-ASR**
(Xiaomi's 8B model, local GPU only, stronger on proper nouns and
code-switching). Both share the same VAD + speaker-clustering stack.
All scripts run directly from the plugin directory — no copying needed.
Define this shorthand at the start of every session:
```bash
SCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/audio-transcribe/scripts
```
## Supported Languages
| `--lang` | Model | Languages | Hotword |
|----------|-------|-----------|---------|
| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |
| `zh-basic` | Paraformer-large | Chinese | No |
| `en` | Paraformer-en | English | No |
| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |
| `whisper` | Whisper-large-v3-turbo | 99_meta.json
{
"ownerId": "kn72agp4n0v3y1gk4qds89r0wn82msn3",
"slug": "zxkane-audio-transcriber-funasr",
"version": "1.7.1",
"publishedAt": 1777649961869
}references/pipeline-details.md
# FunASR Meeting Transcription Pipeline — Technical Details
## Architecture
```
Audio File (.m4a/.mp3/.wav)
│
├─ [ffmpeg] ──► 16kHz mono WAV
│
├─ [Phase 1: FunASR] ──► raw_transcript.json
│ ├─ FSMN-VAD: segment speech vs silence
│ ├─ ASR model (language-dependent, see below)
│ ├─ (Optional) Hotword biasing (SeACo-Paraformer only)
│ ├─ Punctuation restoration (model-dependent)
│ └─ CAM++: speaker embeddings → spectral clustering
│
├─ [Phase 2: Post-process]
│ ├─ Merge consecutive same-speaker utterances (<2s gap)
│ ├─ Map speaker IDs to names (if provided)
│ └─ Auto-verify via self-introduction detection
│
└─ [Phase 3: LLM cleanup] ──► transcript.md
├─ LLM speaker role verification (if --speaker-context provided)
├─ Remove fillers (um, uh, 嗯, 啊, etc.)
├─ Fix ASR errors (homophones, context-based)
├─ Polish grammar while preserving meaning
└─ (Optional) Identify merged speakers via context
```
## Language Presets & Models
### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support
| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |
| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |
| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |
| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |
SeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)
to bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.
### `--lang zh-basic` (Chinese, no hotword)
Same as `zh` but uses the base Paraformer-large without hotword support.
Use when hotword biasing is unnecessary or causing issues with English terms.
| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |
### `--lang en` (English)
| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |
### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)
| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |
SenseVoiceSmall includes built-in punctuation and supports emotion detection.
### `--lang whisper` (Multilingual, 99 languages)
| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |
All presets share the same VAD (`fsmn-vad`) and speaker diariskill-card.md
## Description: Audio Transcribe helps an agent transcribe meeting, podcast, and interview recordings into structured transcripts with ASR, speaker diarization, optional speaker verification, and optional LLM cleanup. This skill is ready for commercial/non-commercial use. ## Publisher: [zxkane](https://clawhub.ai/user/zxkane) ### License/Terms of Use: MIT-0 ## Use Case: Developers, operators, and external users can use this skill to convert audio recordings into Markdown transcripts with speaker labels, timestamps, and optional cleanup for meeting notes or podcast/interview review. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Sensitive audio, transcript excerpts, speaker context, or reference material may be sent to external LLM providers when cleanup is enabled. Mitigation: Omit --model or use --skip-llm to keep cleanup local, and only provide trusted reference and speaker-context files when LLM cleanup is needed. Risk: Speaker-gender inference is enabled by default and may create unnecessary demographic processing. Mitigation: Use --no-detect-gender when demographic inference is not required. Risk: OpenAI-compatible endpoints and setup commands can change the operational trust boundary. Mitigation: Review any OPENAI_BASE_URL value and run setup commands in an isolated virtual environment after reviewing privileged and dependency behavior. ## Reference(s): - [Project homepage](https://github.com/zxkane/audio-transcriber) - [Pipeline Details](references/pipeline-details.md) - [Xiaomi MiMo-V2.5-ASR Model Card](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR) - [ClawHub Skill Page](https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr) ## Skill Output: **Output Type(s):** [text, markdown, JSON, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with inline shell commands; generated transcript output may include Markdown and raw JSON files.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Transcription output can include speaker labels, timestamps, hotword-biased ASR results, and optional LLM-cleaned text.] ## Skill Version(s): 1.7.1 (source: SKILL.md frontmatter and server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr",
"sourceUrl": "https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T01:27:41.989Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T01:27:41.989Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.2K downloads",
"href": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
"sourceUrl": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T01:27:41.989Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.7.1",
"href": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
"sourceUrl": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-01T15:39:21.869Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.7.1",
"description": "**MiMo-V2.5-ASR local GPU transcription and major script rename** - Added support for Xiaomi MiMo-V2.5-ASR (8B, local GPU) with improved proper-noun and code-switching accuracy. - Unified and renamed main script: replaced transcribe_funasr.py with transcribe.py for both FunASR and MiMo workflows. - Project renamed to \"audio-transcribe\" (was \"funasr-transcribe\"); repository and SKILL.md updated accordingly. - README and documentation updated to reflect MiMo engine, requirements (>=20GB GPU VRAM), and combined workflows. - Internal helper/test scripts and pipeline references updated for the new structure.",
"href": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
"sourceUrl": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-01T15:39:21.869Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
