agentCLAWHUBUnverified

Audio Transcribe

This skill should be used when the user explicitly asks to "transcribe a meeting", "transcribe audio", "transcribe a meeting recording", "convert audio to te...

OpenClaw

Rank

62

Safety

84

Downloads

1.2k

Updated

Oct 11, 2026

Version

1.7.1

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.2K downloadsadoption · observed Oct 11, 2026
Latest release
1.7.1release · observed May 1, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr
  1. Install using `clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/snapshot"

Documentation

CLAWHUB

149,258 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: audio-transcribe
version: 1.7.1
description: >
  This skill should be used when the user explicitly asks to "transcribe a meeting",
  "transcribe audio", "transcribe a meeting recording",
  "convert audio to text", "generate meeting minutes from audio",
  "do speech-to-text", "transcribe with speaker diarization",
  "identify speakers in audio", "transcribe Chinese audio",
  "transcribe English audio", "transcribe Japanese audio",
  "multi-speaker transcription", "transcribe a podcast",
  "transcribe podcast episode", "transcribe an interview",
  "convert podcast to text", "podcast to transcript",
  or mentions FunASR, Paraformer, SenseVoice, Whisper, MiMo, MiMo-V2.5-ASR,
  meeting transcription, podcast transcription, or speaker diarization.
  Supports multi-speaker meeting and podcast transcription in Chinese,
  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper),
  plus Xiaomi MiMo-V2.5-ASR (8B, local GPU) for stronger proper-noun and
  code-switching accuracy. Automatic speaker diarization via CAM++,
  hotword biasing (FunASR path), LLM cleanup. FunASR works on GPU and CPU;
  MiMo requires a local CUDA GPU with >=20GB VRAM.
metadata:
  openclaw:
    requires:
      bins: ["python3", "ffmpeg"]
    env_vars:
      - name: AWS_REGION
        required: false
        description: "AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed."
      - name: ANTHROPIC_API_KEY
        required: false
        description: "API key for Anthropic Claude LLM cleanup"
      - name: OPENAI_API_KEY
        required: false
        description: "API key for OpenAI-compatible LLM cleanup"
      - name: OPENAI_BASE_URL
        required: false
        description: "Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)"
    emoji: "🎙️"
    homepage: "https://github.com/zxkane/audio-transcriber"
---

# Meeting & Podcast Transcription (FunASR + MiMo)

Transcribe multi-speaker audio into structured Markdown with automatic
speaker diarization, hotword biasing, and optional LLM cleanup. Two
ASR engine families are available: **FunASR** (Paraformer / SenseVoice /
Whisper — fast, cheap, GPU or CPU, 99 languages) and **MiMo-V2.5-ASR**
(Xiaomi's 8B model, local GPU only, stronger on proper nouns and
code-switching). Both share the same VAD + speaker-clustering stack.

All scripts run directly from the plugin directory — no copying needed.
Define this shorthand at the start of every session:

```bash
SCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/audio-transcribe/scripts
```

## Supported Languages

| `--lang` | Model | Languages | Hotword |
|----------|-------|-----------|---------|
| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |
| `zh-basic` | Paraformer-large | Chinese | No |
| `en` | Paraformer-en | English | No |
| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |
| `whisper` | Whisper-large-v3-turbo | 99

_meta.json

{
  "ownerId": "kn72agp4n0v3y1gk4qds89r0wn82msn3",
  "slug": "zxkane-audio-transcriber-funasr",
  "version": "1.7.1",
  "publishedAt": 1777649961869
}

references/pipeline-details.md

# FunASR Meeting Transcription Pipeline — Technical Details

## Architecture

```
Audio File (.m4a/.mp3/.wav)
  │
  ├─ [ffmpeg] ──► 16kHz mono WAV
  │
  ├─ [Phase 1: FunASR] ──► raw_transcript.json
  │   ├─ FSMN-VAD: segment speech vs silence
  │   ├─ ASR model (language-dependent, see below)
  │   ├─ (Optional) Hotword biasing (SeACo-Paraformer only)
  │   ├─ Punctuation restoration (model-dependent)
  │   └─ CAM++: speaker embeddings → spectral clustering
  │
  ├─ [Phase 2: Post-process]
  │   ├─ Merge consecutive same-speaker utterances (<2s gap)
  │   ├─ Map speaker IDs to names (if provided)
  │   └─ Auto-verify via self-introduction detection
  │
  └─ [Phase 3: LLM cleanup] ──► transcript.md
      ├─ LLM speaker role verification (if --speaker-context provided)
      ├─ Remove fillers (um, uh, 嗯, 啊, etc.)
      ├─ Fix ASR errors (homophones, context-based)
      ├─ Polish grammar while preserving meaning
      └─ (Optional) Identify merged speakers via context
```

## Language Presets & Models

### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support

| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |
| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |
| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |
| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |

SeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)
to bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.

### `--lang zh-basic` (Chinese, no hotword)

Same as `zh` but uses the base Paraformer-large without hotword support.
Use when hotword biasing is unnecessary or causing issues with English terms.

| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |

### `--lang en` (English)

| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |

### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)

| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |

SenseVoiceSmall includes built-in punctuation and supports emotion detection.

### `--lang whisper` (Multilingual, 99 languages)

| Component | Model ID | Params | Role |
|-----------|----------|--------|------|
| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |

All presets share the same VAD (`fsmn-vad`) and speaker diari

skill-card.md

## Description:

Audio Transcribe helps an agent transcribe meeting, podcast, and interview recordings into structured transcripts with ASR, speaker diarization, optional speaker verification, and optional LLM cleanup.

This skill is ready for commercial/non-commercial use.

## Publisher:

[zxkane](https://clawhub.ai/user/zxkane)

### License/Terms of Use:

MIT-0

## Use Case:

Developers, operators, and external users can use this skill to convert audio recordings into Markdown transcripts with speaker labels, timestamps, and optional cleanup for meeting notes or podcast/interview review.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Sensitive audio, transcript excerpts, speaker context, or reference material may be sent to external LLM providers when cleanup is enabled.

Mitigation: Omit --model or use --skip-llm to keep cleanup local, and only provide trusted reference and speaker-context files when LLM cleanup is needed.

Risk: Speaker-gender inference is enabled by default and may create unnecessary demographic processing.

Mitigation: Use --no-detect-gender when demographic inference is not required.

Risk: OpenAI-compatible endpoints and setup commands can change the operational trust boundary.

Mitigation: Review any OPENAI_BASE_URL value and run setup commands in an isolated virtual environment after reviewing privileged and dependency behavior.

## Reference(s):

- [Project homepage](https://github.com/zxkane/audio-transcriber)
- [Pipeline Details](references/pipeline-details.md)
- [Xiaomi MiMo-V2.5-ASR Model Card](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR)
- [ClawHub Skill Page](https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr)

## Skill Output:

**Output Type(s):** [text, markdown, JSON, shell commands, configuration, guidance]

**Output Format:** [Markdown guidance with inline shell commands; generated transcript output may include Markdown and raw JSON files.]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Transcription output can include speaker labels, timestamps, hotword-biased ASR results, and optional LLM-cleaned text.]

## Skill Version(s):

1.7.1 (source: SKILL.md frontmatter and server release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr",
      "sourceUrl": "https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T01:27:41.989Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T01:27:41.989Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.2K downloads",
      "href": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
      "sourceUrl": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T01:27:41.989Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.7.1",
      "href": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
      "sourceUrl": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-01T15:39:21.869Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.7.1",
      "description": "**MiMo-V2.5-ASR local GPU transcription and major script rename** - Added support for Xiaomi MiMo-V2.5-ASR (8B, local GPU) with improved proper-noun and code-switching accuracy. - Unified and renamed main script: replaced transcribe_funasr.py with transcribe.py for both FunASR and MiMo workflows. - Project renamed to \"audio-transcribe\" (was \"funasr-transcribe\"); repository and SKILL.md updated accordingly. - README and documentation updated to reflect MiMo engine, requirements (>=20GB GPU VRAM), and combined workflows. - Internal helper/test scripts and pipeline references updated for the new structure.",
      "href": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
      "sourceUrl": "https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-01T15:39:21.869Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to Audio Transcribe and adjacent AI workflows.