{"id":"1bca6a0f-3f27-41ff-890f-8c6f0339706c","entityType":"agent","slug":"clawhub-zxkane-zxkane-audio-transcriber-funasr","name":"Audio Transcribe","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zxkane-zxkane-audio-transcriber-funasr","canonicalPath":"/agent/clawhub-zxkane-zxkane-audio-transcriber-funasr","generatedAt":"2026-10-11T03:57:30.131Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T01:27:41.989Z","emptyReason":null},"description":"This skill should be used when the user explicitly asks to \"transcribe a meeting\", \"transcribe audio\", \"transcribe a meeting recording\", \"convert audio to te...","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr","sourceUrl":"https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr","homepage":"https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":62,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Audio Transcribe technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T01:27:41.989Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T01:27:41.989Z","emptyReason":null},"stars":null,"forks":null,"downloads":1206,"packageName":null,"latestVersion":"1.7.1","tractionLabel":"1.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T01:27:41.919Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T01:27:41.989Z","lastCrawledAt":"2026-10-11T01:27:41.919Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T01:27:41.919Z","lastVerifiedAt":null,"highlights":[{"version":"1.7.1","createdAt":"2026-05-01T15:39:21.869Z","changelog":"**MiMo-V2.5-ASR local GPU transcription and major script rename** - Added support for Xiaomi MiMo-V2.5-ASR (8B, local GPU) with improved proper-noun and code-switching accuracy. - Unified and renamed main script: replaced transcribe_funasr.py with transcribe.py for both FunASR and MiMo workflows. - Project renamed to \"audio-transcribe\" (was \"funasr-transcribe\"); repository and SKILL.md updated accordingly. - README and documentation updated to reflect MiMo engine, requirements (>=20GB GPU VRAM), and combined workflows. - Internal helper/test scripts and pipeline references updated for the new structure.","fileCount":14,"zipByteSize":87017},{"version":"1.7.0","createdAt":"2026-05-01T00:42:05.731Z","changelog":"MiMo V2.5 ASR and `mimo` mode added; LLM provider/config options improved. - Added support for MiMo V2.5 ASR (`--lang mimo`) for faster local transcription (GPU only, Chinese/English/dialects/code-switch). - Updated quick start and language table to include MiMo and its diarization details. - Improved LLM cleanup: now requires explicit `--provider` (bedrock, anthropic, openai) for clarity and correct routing. - Updated environment/setup scripts to install MiMo dependencies and support new model. - Added scripts for MiMo (`mimo_asr.py`, `setup_mimo.sh`, `test_mimo_asr.py`). - Improved docs with more details on Bedrock/global and OpenAI/Anthropic LLM routing.","fileCount":13,"zipByteSize":85490},{"version":"1.6.0","createdAt":"2026-04-28T11:06:41.390Z","changelog":"- Added speaker gender detection script (`scripts/speaker_gender.py`). - Improved main transcription and speaker verification scripts. - Documentation updated to reflect new features and changes.","fileCount":10,"zipByteSize":60179},{"version":"1.5.1","createdAt":"2026-04-26T01:11:31.256Z","changelog":"Version 1.5.1 - Improved logic in transcribe_funasr.py for handling speaker verification and LLM cleanup stages. - Enhanced speaker label verification in scripts/test_speaker_verification.py, including better mismatch detection and reporting. - Updated documentation in SKILL.md to clarify workflow and usage, with small corrections and enhanced guidance.","fileCount":9,"zipByteSize":52135},{"version":"1.5.0","createdAt":"2026-04-25T14:23:21.598Z","changelog":"Version 1.5.0 - Adds a clear warning and detection logic for incorrect use of podcast aliases in `--speakers`, prompting for the real host/guest name based on supporting materials. - Updates SKILL.md documentation to explain the new requirement: `--speakers` must use actual real names, not show or podcast aliases. - Guidance added for including both real names and aliases in `hotwords.txt` for better ASR recognition. - No changes to skill APIs or setup; this update improves correctness of speaker labeling and overall transcript quality.","fileCount":9,"zipByteSize":51682},{"version":"1.4.1","createdAt":"2026-04-25T07:20:06.124Z","changelog":"# zxkane-audio-transcriber-funasr 1.4.1 - Added a new `test_speaker_verification.py` script for improved speaker verification testing. - Updated documentation in SKILL.md to match current features and clarify supported workflows. - Minor tweaks and improvements to `transcribe_funasr.py` to enhance transcription reliability and maintain consistency with speaker verification utilities.","fileCount":9,"zipByteSize":49924},{"version":"1.4.0","createdAt":"2026-04-21T05:16:00.060Z","changelog":"- LLM cleanup (Phase 3) is now opt-in and runs only when --model is specified; by default, all transcription is local and no data is sent to external services. - Documentation clarifies privacy defaults: external LLMs are used only with --model, and Bedrock uses the AWS credential chain. - Description updated to reflect that transcription triggers only when the user explicitly requests it. - Example commands and workflow sections updated to reflect the new LLM cleanup activation method. - No major behavioral or interface changes beyond LLM phase invocation and stricter privacy defaults.","fileCount":9,"zipByteSize":46692},{"version":"1.3.1","createdAt":"2026-04-21T03:13:17.389Z","changelog":"**Adds environment variable docs and improves setup automation.** - Documents support for Anthropic, OpenAI-compatible, and Bedrock LLM providers via environment variables (ANTHROPIC_API_KEY, OPENAI_API_KEY, AWS_REGION, etc.) - Updates SKILL.md to explain LLM provider credentials, API base URL, and privacy/credential isolation - Adds \"AUTO_YES=1\" usage to environment setup for non-interactive installs - No behavioral changes to transcription logic itself","fileCount":9,"zipByteSize":46642}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17dzqxt6ty0dfybdr3f7wh31h852ebv:zxkane-audio-transcriber-funasr` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zxkane/zxkane-audio-transcriber-funasr before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T03:57:30.127Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zxkane-zxkane-audio-transcriber-funasr/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T01:27:41.989Z","emptyReason":null},"readme":"Skill: Audio Transcribe\n\nOwner: zxkane\n\nSummary: This skill should be used when the user explicitly asks to \"transcribe a meeting\", \"transcribe audio\", \"transcribe a meeting recording\", \"convert audio to te...\n\nTags: latest:1.7.1\n\nVersion history:\n\nv1.7.1 | 2026-05-01T15:39:21.869Z | auto\n\n**MiMo-V2.5-ASR local GPU transcription and major script rename**\n\n- Added support for Xiaomi MiMo-V2.5-ASR (8B, local GPU) with improved proper-noun and code-switching accuracy.\n- Unified and renamed main script: replaced transcribe_funasr.py with transcribe.py for both FunASR and MiMo workflows.\n- Project renamed to \"audio-transcribe\" (was \"funasr-transcribe\"); repository and SKILL.md updated accordingly.\n- README and documentation updated to reflect MiMo engine, requirements (>=20GB GPU VRAM), and combined workflows.\n- Internal helper/test scripts and pipeline references updated for the new structure.\n\nv1.7.0 | 2026-05-01T00:42:05.731Z | auto\n\nMiMo V2.5 ASR and `mimo` mode added; LLM provider/config options improved.\n\n- Added support for MiMo V2.5 ASR (`--lang mimo`) for faster local transcription (GPU only, Chinese/English/dialects/code-switch).\n- Updated quick start and language table to include MiMo and its diarization details.\n- Improved LLM cleanup: now requires explicit `--provider` (bedrock, anthropic, openai) for clarity and correct routing.\n- Updated environment/setup scripts to install MiMo dependencies and support new model.\n- Added scripts for MiMo (`mimo_asr.py`, `setup_mimo.sh`, `test_mimo_asr.py`).\n- Improved docs with more details on Bedrock/global and OpenAI/Anthropic LLM routing.\n\nv1.6.0 | 2026-04-28T11:06:41.390Z | auto\n\n- Added speaker gender detection script (`scripts/speaker_gender.py`).\n- Improved main transcription and speaker verification scripts.\n- Documentation updated to reflect new features and changes.\n\nv1.5.1 | 2026-04-26T01:11:31.256Z | auto\n\nVersion 1.5.1\n\n- Improved logic in transcribe_funasr.py for handling speaker verification and LLM cleanup stages.\n- Enhanced speaker label verification in scripts/test_speaker_verification.py, including better mismatch detection and reporting.\n- Updated documentation in SKILL.md to clarify workflow and usage, with small corrections and enhanced guidance.\n\nv1.5.0 | 2026-04-25T14:23:21.598Z | auto\n\nVersion 1.5.0\n\n- Adds a clear warning and detection logic for incorrect use of podcast aliases in `--speakers`, prompting for the real host/guest name based on supporting materials.\n- Updates SKILL.md documentation to explain the new requirement: `--speakers` must use actual real names, not show or podcast aliases.\n- Guidance added for including both real names and aliases in `hotwords.txt` for better ASR recognition.\n- No changes to skill APIs or setup; this update improves correctness of speaker labeling and overall transcript quality.\n\nv1.4.1 | 2026-04-25T07:20:06.124Z | auto\n\n# zxkane-audio-transcriber-funasr 1.4.1\n\n- Added a new `test_speaker_verification.py` script for improved speaker verification testing.\n- Updated documentation in SKILL.md to match current features and clarify supported workflows.\n- Minor tweaks and improvements to `transcribe_funasr.py` to enhance transcription reliability and maintain consistency with speaker verification utilities.\n\nv1.4.0 | 2026-04-21T05:16:00.060Z | auto\n\n- LLM cleanup (Phase 3) is now opt-in and runs only when --model is specified; by default, all transcription is local and no data is sent to external services.\n- Documentation clarifies privacy defaults: external LLMs are used only with --model, and Bedrock uses the AWS credential chain.\n- Description updated to reflect that transcription triggers only when the user explicitly requests it.\n- Example commands and workflow sections updated to reflect the new LLM cleanup activation method.\n- No major behavioral or interface changes beyond LLM phase invocation and stricter privacy defaults.\n\nv1.3.1 | 2026-04-21T03:13:17.389Z | auto\n\n**Adds environment variable docs and improves setup automation.**\n\n- Documents support for Anthropic, OpenAI-compatible, and Bedrock LLM providers via environment variables (ANTHROPIC_API_KEY, OPENAI_API_KEY, AWS_REGION, etc.)\n- Updates SKILL.md to explain LLM provider credentials, API base URL, and privacy/credential isolation\n- Adds \"AUTO_YES=1\" usage to environment setup for non-interactive installs\n- No behavioral changes to transcription logic itself\n\nv1.3.0 | 2026-04-20T21:25:28.498Z | auto\n\n- Expanded language support: now handles Chinese, English, Japanese, Korean, Cantonese, and 99 languages (via Whisper), with automatic speaker diarization and hotword biasing.\n- New, detailed workflow: guides users to provide context like meeting type, participant names, supporting documents, and preference for language and number of speakers to optimize transcription quality.\n- Enhanced presets and diarization: per-language model selection with clear caveats on diarization support, especially for `auto` and `whisper` modes.\n- LLM optional cleanup: supports post-processing transcripts with Bedrock, Anthropic, or OpenAI-compatible LLMs, with resume and skip options.\n- Utility scripts included: speaker verification and reassignment script helps detect and fix swapped or misidentified speakers.\n- Audio preprocessing improvements: all inputs auto-converted to 16kHz mono FLAC for reliability, with detailed format recommendations.\n\nArchive index:\n\nArchive v1.7.1: 14 files, 87017 bytes\n\nFiles: references/pipeline-details.md (20541b), scripts/llm_utils.py (7469b), scripts/mimo_asr.py (19140b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (5893b), scripts/setup_mimo.sh (6125b), scripts/speaker_gender.py (11457b), scripts/test_mimo_asr.py (24564b), scripts/test_speaker_verification.py (92315b), scripts/transcribe.py (68727b), scripts/verify_speakers.py (18210b), skill-card.md (2513b), SKILL.md (15771b), _meta.json (150b)\n\nFile v1.7.1:SKILL.md\n\n---\nname: audio-transcribe\nversion: 1.7.1\ndescription: >\n  This skill should be used when the user explicitly asks to \"transcribe a meeting\",\n  \"transcribe audio\", \"transcribe a meeting recording\",\n  \"convert audio to text\", \"generate meeting minutes from audio\",\n  \"do speech-to-text\", \"transcribe with speaker diarization\",\n  \"identify speakers in audio\", \"transcribe Chinese audio\",\n  \"transcribe English audio\", \"transcribe Japanese audio\",\n  \"multi-speaker transcription\", \"transcribe a podcast\",\n  \"transcribe podcast episode\", \"transcribe an interview\",\n  \"convert podcast to text\", \"podcast to transcript\",\n  or mentions FunASR, Paraformer, SenseVoice, Whisper, MiMo, MiMo-V2.5-ASR,\n  meeting transcription, podcast transcription, or speaker diarization.\n  Supports multi-speaker meeting and podcast transcription in Chinese,\n  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper),\n  plus Xiaomi MiMo-V2.5-ASR (8B, local GPU) for stronger proper-noun and\n  code-switching accuracy. Automatic speaker diarization via CAM++,\n  hotword biasing (FunASR path), LLM cleanup. FunASR works on GPU and CPU;\n  MiMo requires a local CUDA GPU with >=20GB VRAM.\nmetadata:\n  openclaw:\n    requires:\n      bins: [\"python3\", \"ffmpeg\"]\n    env_vars:\n      - name: AWS_REGION\n        required: false\n        description: \"AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed.\"\n      - name: ANTHROPIC_API_KEY\n        required: false\n        description: \"API key for Anthropic Claude LLM cleanup\"\n      - name: OPENAI_API_KEY\n        required: false\n        description: \"API key for OpenAI-compatible LLM cleanup\"\n      - name: OPENAI_BASE_URL\n        required: false\n        description: \"Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)\"\n    emoji: \"🎙️\"\n    homepage: \"https://github.com/zxkane/audio-transcriber\"\n---\n\n# Meeting & Podcast Transcription (FunASR + MiMo)\n\nTranscribe multi-speaker audio into structured Markdown with automatic\nspeaker diarization, hotword biasing, and optional LLM cleanup. Two\nASR engine families are available: **FunASR** (Paraformer / SenseVoice /\nWhisper — fast, cheap, GPU or CPU, 99 languages) and **MiMo-V2.5-ASR**\n(Xiaomi's 8B model, local GPU only, stronger on proper nouns and\ncode-switching). Both share the same VAD + speaker-clustering stack.\n\nAll scripts run directly from the plugin directory — no copying needed.\nDefine this shorthand at the start of every session:\n\n```bash\nSCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/audio-transcribe/scripts\n```\n\n## Supported Languages\n\n| `--lang` | Model | Languages | Hotword |\n|----------|-------|-----------|---------|\n| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |\n| `zh-basic` | Paraformer-large | Chinese | No |\n| `en` | Paraformer-en | English | No |\n| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |\n| `whisper` | Whisper-large-v3-turbo | 99 languages | No |\n| `mimo` | MiMo-V2.5-ASR (local 8B, GPU-only) | zh/en/code-switch/dialects | No |\n\nAll presets include **speaker diarization** (CAM++) and **VAD** (FSMN).\n`mimo` reuses the FSMN VAD + CAM++ stack around MiMo's text output.\n\n> **Diarization caveat:** `auto` and `whisper` do not output per-sentence timestamps,\n> so speaker diarization does not work with these presets. Use `zh`, `zh-basic`,\n> `en`, or `mimo` when speaker identification is needed (e.g., podcasts, meetings).\n\n## Workflow\n\nBefore starting transcription, **always ask the user**:\n\n1. **Audio file** — path to the recording (required)\n2. **Type** — meeting, podcast, or interview? (affects defaults)\n3. **Language** — what language is spoken? (default: Chinese)\n4. **Number of speakers** — how many participants? (improves diarization)\n5. **Speaker names** — for podcasts: host + guest names; for meetings: attendee list\n6. **Supporting files** — ask:\n   > \"Do you have any of the following to improve accuracy?\"\n   > - **Attendee / guest list** — for hotwords and speaker mapping\n   > - **Meeting agenda or episode topic** — for hotwords (terms, names)\n   > - **Reference documents** (show notes, prior notes) — for speaker identification and ASR correction\n\n**Adapt defaults by recording type:**\n- **Meeting**: default `--lang zh`, ask about supporting files\n- **Podcast / interview**: default `--lang zh`, `--num-speakers 2`, always ask for\n  host + guest names, suggest `--speaker-context` for roles\n  (do NOT use `--lang auto` — it lacks timestamps for speaker diarization)\n\n> **⚠️ `--speakers` must use the speaker's real name, not a podcast alias.**\n> The value passed to `--speakers` is used verbatim as the speaker label in the\n> output transcript. Always derive it from the host/guest's actual name (e.g.\n> from a shownotes \"Host:\" field), not from the podcast feed name or title.\n>\n> Example: if shownotes lists \"Host: 张三（张三的播客）\", pass `--speakers '张三'`\n> — not the alias \"张三的播客\". Add both the real name and the alias to\n> `hotwords.txt` so ASR can recognise both forms.\n>\n> When both `--speakers` and `--reference` are supplied, the script detects\n> this mistake at startup and prints an `ACTION REQUIRED` block naming the\n> suggested real name. **If you see that block, stop the run and re-invoke\n> with the corrected `--speakers` value before Phase 3** — the warning does\n> not abort the pipeline.\n\nIf the user provides supporting materials:\n- Extract participant names and key terms → create `hotwords.txt` (include both real name and alias)\n- Extract per-person context → create `speaker-context.json`\n- Pass original reference document with `--reference`\n- Use all three together for best results\n\n## Quick Start\n\n### 1. Environment Setup\n\n```bash\nAUTO_YES=1 bash $SCRIPTS/setup_env.sh\n# Or force CPU:  AUTO_YES=1 bash $SCRIPTS/setup_env.sh cpu\n```\n\nThe setup script patches FunASR's spectral clustering for O(N²·k) performance.\nWithout this, recordings over ~1 hour hang for hours during speaker clustering.\n\n### 2. Run Transcription\n\nOutput files are written to the current working directory.\n\n**LLM cleanup (Phase 3) is opt-in.** By default, transcription runs locally\nwithout contacting any external service. To enable LLM-powered ASR correction\nand speaker name refinement, pass `--model <model-id>`. Use LLM cleanup when:\n- The raw transcript has many ASR errors (names, technical terms)\n- You need polished, publication-ready output\n- Speaker names need to be refined from context\n\n> **⚠️ Data Privacy:** When LLM cleanup is enabled via `--model`, transcript\n> excerpts are sent to external LLM providers (AWS Bedrock, Anthropic, or\n> OpenAI depending on the model ID). Use `--skip-llm` or omit `--model` to\n> keep all data local. For Bedrock, boto3 uses the standard AWS credential\n> chain (IAM role, SSO, `~/.aws/credentials`, env vars).\n\n```bash\n# Chinese meeting with hotwords (local-only, no LLM)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt\n\n# English meeting with speaker names\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang en --speakers \"Alice,Bob,Carol,Dave\"\n\n# Auto-detect language (zh/en/ja/ko/yue)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang auto --num-speakers 6\n\n# Whisper for any language\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang whisper --num-speakers 4\n\n# Enable LLM cleanup for polished output (requires --model)\n# Bedrock (uses AWS credential chain: IAM role, SSO, ~/.aws/credentials)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt \\\n    --provider bedrock --model us.anthropic.claude-sonnet-4-6\n\n# Bedrock \"global\" cross-region profile (recent AWS deployments)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider bedrock --model global.anthropic.claude-sonnet-4-6\n\n# Bedrock via litellm-style wrapper (supported; prefix is stripped for boto3)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider bedrock --model amazon-bedrock/global.anthropic.claude-sonnet-4-6\n\n# Anthropic API (requires ANTHROPIC_API_KEY env var)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider anthropic --model claude-sonnet-4-6\n\n# OpenAI-compatible API (requires OPENAI_API_KEY env var)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider openai --model gpt-4o\n\n# Full pipeline with all supporting files + LLM (best quality)\npython3 $SCRIPTS/transcribe.py episode.m4a \\\n    --lang zh --num-speakers 2 \\\n    --hotwords hotwords.txt \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json \\\n    --reference show-notes.md \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Resume interrupted LLM cleanup\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --skip-transcribe --model us.anthropic.claude-sonnet-4-6\n```\n\n### 3. Verify Speaker Labels\n\nIf the transcript has swapped speaker labels (common with podcasts),\nthe verification script can detect and fix mismatches using LLM analysis:\n\n```bash\n# Dry-run: check if host/guest are swapped\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json\n\n# Apply the fix\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json --fix\n\n# Multi-speaker meeting: full reassignment\npython3 $SCRIPTS/verify_speakers.py meeting_raw_transcript.json \\\n    --speakers \"Alice,Bob,Carol,Dave\" \\\n    --speaker-context speaker-context.json --fix\n\n# Then regenerate the markdown with corrected labels\npython3 $SCRIPTS/transcribe.py original.m4a \\\n    --skip-transcribe --clean-cache\n```\n\nThe script analyzes the first 5 minutes (configurable with `--minutes`)\nand auto-detects podcast (2 speakers, swap detection) vs meeting\n(N speakers, full reassignment).\n\n## Audio Preprocessing\n\nThe script automatically converts input audio to 16kHz mono FLAC and\nvalidates that no audio is lost (detects silent truncation).\n\n| Format | 4h14m meeting | Quality | Recommendation |\n|--------|--------------|---------|----------------|\n| **FLAC** | **219MB** | Lossless | **Default, safest** |\n| Opus | 55MB | Lossy | Risk of truncation on long files |\n| WAV | 465MB | Lossless | Works but larger |\n| Original M4A | 173MB | Source | Also works directly |\n\n**Do NOT split long recordings** — splitting breaks speaker ID consistency.\n\n## MiMo-V2.5-ASR (optional, GPU-only)\n\n`--lang mimo` runs Xiaomi's\n[MiMo-V2.5-ASR](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR) locally on a\nCUDA GPU. Use it when:\n- You want to evaluate MiMo against Paraformer on Chinese audio.\n- The recording has heavy code-switching, dialects (Wu, Cantonese, Hokkien,\n  Sichuanese), lyrics, or rare proper nouns that other presets mis-transcribe.\n\n**Requirements:**\n- CUDA ≥12.0 and **≥20 GB VRAM** (16 GB cards OOM during inference).\n- Python 3.12 (enforced by `setup_env.sh`).\n- ~20 GB weight download (one-time) and `flash-attn==2.7.4.post1` compile\n  (needs `nvcc` from the CUDA toolkit, takes 10–30 min).\n\n**Install (opt-in):**\n\n```bash\n# One-time: install MiMo on top of the standard environment\nAUTO_YES=1 INSTALL_MIMO=1 \\\n    MIMO_WEIGHTS_PATH=/mnt/models/hf \\\n    bash $SCRIPTS/setup_env.sh\n```\n\n**Run:**\n\n```bash\npython3 $SCRIPTS/transcribe.py podcast.m4a \\\n    --lang mimo --num-speakers 2 \\\n    --mimo-weights-path /mnt/models/hf\n```\n\n**Resume after failure:**\n\n```bash\npython3 $SCRIPTS/transcribe.py podcast.m4a \\\n    --lang mimo --resume-mimo --mimo-weights-path /mnt/models/hf\n```\n\n**Limitations:**\n- No hotword biasing (MiMo has no API for it — `--hotwords` is ignored).\n- No CPU fallback.\n- Inference is slower than Paraformer on the same GPU (8B model vs ~0.3B);\n  expect RTF around 0.1–0.2 on an A100.\n\n## Key Flags\n\n| Flag | Purpose |\n|------|---------|\n| `--lang` | `zh` (default), `zh-basic`, `en`, `auto`, `whisper` |\n| `--hotwords` | Hotword file or string — biases ASR (zh only) |\n| `--reference F` | Reference file for LLM ASR correction |\n| `--num-speakers N` | Expected speaker count (improves diarization) |\n| `--speakers \"A,B,C\"` | Assign real names by first-appearance order |\n| `--speaker-context F` | JSON with per-speaker roles for LLM |\n| `--no-detect-gender` | Disable automatic speaker gender detection (CAM++ gender classifier) |\n| `--speaker-genders \"A:female,B:male\"` | Override per-speaker gender (also accepts positional `female,male`) |\n| `--audio-format` | `flac` (default), `opus`, `wav` |\n| `--device cpu` | Force CPU mode |\n| `--batch-size N` | Adjust for memory (60 for CPU, 100 if GPU OOM) |\n| `--phase1-only` | Exit after Phase 1 (VAD + ASR + diarization), skip Phase 2 + 3 |\n| `--json-out PATH` | Write raw transcript JSON to explicit path (overrides default naming) |\n| `--skip-transcribe` | Resume from saved `*_raw_transcript.json` |\n| `--skip-llm` | Skip LLM cleanup (default when `--model` is omitted) |\n| `--model ID` | Enable LLM cleanup with this model (auto-detects Bedrock/Anthropic/OpenAI) |\n| `--title \"...\"` | Output document title |\n| `--clean-cache` | Delete LLM chunk cache after completion |\n| `--output PATH` | Custom output file path |\n| `--model-cache-dir` | ModelScope model cache directory (~3GB, default: `~/.cache/modelscope/`) |\n| `--mimo-audio-tag` | MiMo language hint: `<chinese>` (default), `<english>`, `<auto>` |\n| `--mimo-batch N` | Concurrent VAD segments per MiMo call (default 1; H100/80GB can go higher) |\n| `--mimo-weights-path DIR` | Cache dir for MiMo weights (default: `$HF_HOME` → `~/.cache/huggingface`) |\n| `--resume-mimo` | Resume MiMo Phase 1 from `*_mimo_partial.json` after a mid-run failure |\n\n## Outputs\n\n- `<stem>-transcript.md` — Final Markdown with speaker labels and timestamps\n- `<stem>_raw_transcript.json` — Raw Phase 1 output (for resume/analysis)\n\n## Speaker Diarization Tips\n\nFunASR's CAM++ may merge acoustically similar speakers. To improve:\n\n1. **`--num-speakers N`** — Hint expected count\n2. **`--hotwords`** — Include participant names (Chinese names work best)\n3. **`--speaker-context`** — Provide per-person keywords for LLM splitting\n4. **Keyword matching** — Search `*_raw_transcript.json` for unique phrases\n\n### Speaker gender\n\nEnabled by default: each detected speaker is classified as `male` / `female`\nvia 3D-Speaker's CAM++ gender classifier (`iic/speech_campplus_two_class_gender_16k`).\nThe result appears next to each name in the **Speaker List** table and is\ninjected into the LLM cleanup prompt so pronouns (他/她, he/she) get corrected.\n\nPrecedence when combined:\n\n1. `--speaker-genders \"Alice:female,Bob:male\"` (explicit CLI) — always wins\n2. Reference text hints like `主播（女）：韩梅梅` or `Host (male): Alice` — override auto\n3. CAM++ auto-detection — fallback\n\nDisable with `--no-detect-gender` if you don't need gender and want to save\nthe ~500 MB model download and extra inference time.\n\n## CPU-only / Low-Memory Machines\n\nLong recordings on resource-constrained machines may hit exec timeouts\nor OOM kills. See `references/pipeline-details.md` for workarounds:\n- Detach from agent timeouts with `systemd-run` or `nohup`\n- Prevent OOM via swap and/or `--lang zh-basic` (lighter model)\n\n## Additional Resources\n\n- **`references/pipeline-details.md`** — Architecture, model specs, benchmarks,\n  speaker role verification, hotword effectiveness, clustering patch\n- **`scripts/transcribe.py`** — Main transcription pipeline\n- **`scripts/verify_speakers.py`** — Speaker label verification & fix\n- **`scripts/llm_utils.py`** — Shared LLM infrastructure (Bedrock/Anthropic/OpenAI)\n- **`scripts/setup_env.sh`** — Environment setup (venv + deps + patch)\n\nFile v1.7.1:_meta.json\n\n{\n  \"ownerId\": \"kn72agp4n0v3y1gk4qds89r0wn82msn3\",\n  \"slug\": \"zxkane-audio-transcriber-funasr\",\n  \"version\": \"1.7.1\",\n  \"publishedAt\": 1777649961869\n}\n\nFile v1.7.1:references/pipeline-details.md\n\n# FunASR Meeting Transcription Pipeline — Technical Details\n\n## Architecture\n\n```\nAudio File (.m4a/.mp3/.wav)\n  │\n  ├─ [ffmpeg] ──► 16kHz mono WAV\n  │\n  ├─ [Phase 1: FunASR] ──► raw_transcript.json\n  │   ├─ FSMN-VAD: segment speech vs silence\n  │   ├─ ASR model (language-dependent, see below)\n  │   ├─ (Optional) Hotword biasing (SeACo-Paraformer only)\n  │   ├─ Punctuation restoration (model-dependent)\n  │   └─ CAM++: speaker embeddings → spectral clustering\n  │\n  ├─ [Phase 2: Post-process]\n  │   ├─ Merge consecutive same-speaker utterances (<2s gap)\n  │   ├─ Map speaker IDs to names (if provided)\n  │   └─ Auto-verify via self-introduction detection\n  │\n  └─ [Phase 3: LLM cleanup] ──► transcript.md\n      ├─ LLM speaker role verification (if --speaker-context provided)\n      ├─ Remove fillers (um, uh, 嗯, 啊, etc.)\n      ├─ Fix ASR errors (homophones, context-based)\n      ├─ Polish grammar while preserving meaning\n      └─ (Optional) Identify merged speakers via context\n```\n\n## Language Presets & Models\n\n### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |\n| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |\n| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |\n| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |\n\nSeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)\nto bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.\n\n### `--lang zh-basic` (Chinese, no hotword)\n\nSame as `zh` but uses the base Paraformer-large without hotword support.\nUse when hotword biasing is unnecessary or causing issues with English terms.\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |\n\n### `--lang en` (English)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |\n\n### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |\n\nSenseVoiceSmall includes built-in punctuation and supports emotion detection.\n\n### `--lang whisper` (Multilingual, 99 languages)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |\n\nAll presets share the same VAD (`fsmn-vad`) and speaker diarization (`cam++`) models.\nModels are auto-downloaded from ModelScope on first run.\n\n## Hotword Biasing\n\nSeACo-Paraformer (`--lang zh`) supports hotword customization to improve recognition\nof specific terms — particularly useful for participant names, project names, and\ndomain-specific jargon in meetings.\n\n### How to provide hotwords\n\n```bash\n# Space-separated string\npython3 transcribe.py meeting.wav --lang zh --hotwords \"张三 李四 ClawCon Rebase\"\n\n# Text file (one word per line)\npython3 transcribe.py meeting.wav --lang zh --hotwords hotwords.txt\n```\n\n### What to include in hotwords\n\nFor meeting transcription, a good hotwords file includes:\n- **Participant names** (full names in the meeting's language)\n- **Project / product names** (internal codenames, product brands)\n- **Domain-specific Chinese terms** (technical jargon, acronyms in Chinese)\n- **Organization names** (company, team, department names)\n\n### Effectiveness — empirical results\n\nTested on a 4h14m, 9-speaker Chinese meeting with 27 hotwords:\n\n| Term | Without hotwords | With hotwords | Change |\n|------|-----------------|---------------|--------|\n| 龙虾 (lobechat) | 28 | 42 | **+50%** |\n| 高琦 (person name) | 0 | 7 | **0 → 7** |\n| 搬瓦工 (BandwagonHost) | 0 | 1 | **0 → 1** |\n| 谢锐 (person name) | 0 | 1 | improved |\n| 鲲鹏 (org name) | 6 | 7 | slight improvement |\n| Rebase (English) | 5 | 0 | **regression** |\n| Tailwind (English) | 3 | 1 | **regression** |\n\n**Key findings:**\n1. **Chinese terms benefit significantly** — names, brands, and Chinese jargon\n   see clear improvement (龙虾 +50%, 高琦 from zero)\n2. **English terms may regress** — SeACo's hotword biasing operates on Chinese\n   token vocabulary; English loanwords can be disrupted\n3. **Person names have limited uplift** — meeting participants rarely say full\n   names aloud; hotwords help only when names do appear in speech\n\n**Recommendation:** Include Chinese terms and names in hotwords. For English\ntechnical terms, rely on Phase 3 LLM cleanup rather than hotword biasing.\nIf English term accuracy is critical and hotword biasing causes regressions,\nuse `--lang zh-basic` instead.\n\n## Performance Benchmarks\n\nTested on a 4h14m, 9-speaker Chinese meeting recording (GPU: L40S 46GB):\n\n| Metric | Paraformer-large | SeACo + Hotword | Notes |\n|--------|-----------------|-----------------|-------|\n| Model load | 14s | 14s (422s first run, downloading 944MB model) | Cached after first run |\n| Transcription | 169s | 168s | Virtually identical |\n| Raw sentences | 6672 | 6725 | Comparable |\n| Merged segments | 1695 | 1724 | Comparable |\n| Speakers detected | 7 (of 9) | 7 (of 9) | Same diarization result |\n\n| Metric | GPU (L40S 46GB) | CPU (estimated) |\n|--------|-----------------|-----------------|\n| Model load | 14s | ~30s |\n| Transcription | 169s | ~30-60 min |\n| Speaker clustering | ~10s (patched) | ~2-5 min (patched) |\n| LLM cleanup (17 chunks) | ~35 min | ~35 min (network-bound) |\n| Total | ~38 min | ~70-100 min |\n\n**Without the clustering patch**, the original `scipy.linalg.eigh()` on the full Laplacian\nmatrix was O(N^3) and took **10+ hours** on this recording. The patch reduces it to O(N^2*k)\nvia `scipy.sparse.linalg.eigsh()`.\n\n## Clustering Patch (Critical for Long Meetings)\n\nFunASR's `SpectralCluster.get_spec_embs()` uses `scipy.linalg.eigh(L)` which computes\nALL eigenvalues of the NxN Laplacian. For a 4-hour recording, N can be 6000+, making\nthis O(N^3) operation take hours.\n\nThe patch (`scripts/patch_clustering.py`) replaces this with:\n- `scipy.sparse.linalg.eigsh(L_sparse, k=num_speakers, which='SM')` — only computes\n  the k smallest eigenvalues needed, reducing complexity to O(N^2 * k)\n- Vectorized `p_pruning()` — replaces Python loop with numpy broadcasting\n\n**Always run the patch before processing meetings longer than ~1 hour.**\n\n## Speaker Role Verification\n\nSpeaker names are assigned by first-appearance order in the audio, which may\nswap host/guest labels (especially in podcasts). Two layers of automatic\nverification, plus a standalone post-hoc tool:\n\n1. **Phase 2 — self-introduction detection**: scans the first 5 minutes for\n   explicit self-introductions (\"我是X\", \"I'm X\") and swaps labels if mismatched\n2. **Phase 3 — LLM role verification**: when `--speaker-context` is provided,\n   the LLM analyzes the first chunk (up to 15 minutes) before cleanup begins.\n   For 2 speakers: binary CORRECT/SWAP detection. For 3+ speakers: full\n   JSON-based reassignment matching each label to the correct person.\n3. **Post-hoc — `verify_speakers.py`**: standalone script that verifies any\n   existing `*_raw_transcript.json`. Same two modes (2-speaker swap, N-speaker\n   reassignment) with dry-run support. See SKILL.md § Verify Speaker Labels.\n\nFor podcasts, always provide `--speaker-context` describing host/guest roles.\n\n## Speaker Diarization Limitations\n\nFunASR's CAM++ speaker diarization may merge acoustically similar speakers into one ID.\nIn the tested 9-person meeting, only 7 unique IDs were detected (two pairs merged).\n\nWorkarounds:\n1. **Provide `--num-speakers N`** to hint expected count (uses `preset_spk_num`)\n2. **Post-hoc keyword matching**: use reference documents (meeting agendas, attendee notes)\n   to identify which speaker ID maps to which person\n3. **LLM-assisted splitting**: provide `--speaker-context` with per-person keywords;\n   the LLM can then split merged speakers when context is clear (~73% success rate)\n\n## Supporting Files for Better Results\n\nPrepare these files before transcription for best results:\n\n| File | Used in | Purpose |\n|------|---------|---------|\n| `hotwords.txt` | Phase 1 (`--hotwords`) | Bias ASR toward names and terms |\n| `speaker-context.json` | Phase 3 (`--speaker-context`) | Help LLM identify and split speakers |\n| Meeting agenda | Manual reference | Identify meeting phases for post-analysis |\n| Attendee list | Build hotwords + speaker names | Map speaker IDs to real names |\n\n### Example: preparing supporting files from a meeting invite\n\n```bash\n# 1. Create hotwords.txt from attendee list and agenda\ncat > hotwords.txt << 'EOF'\nAlice\nBob\nCarol\nProjectAlpha\nSprint Review\nQ2 OKR\nEOF\n\n# 2. Create speaker-context.json from attendee roles\ncat > speaker-context.json << 'EOF'\n{\n  \"Alice\": \"Engineering manager, discusses sprint velocity and tech debt\",\n  \"Bob\": \"Product manager, presents roadmap and customer feedback\",\n  \"Carol\": \"Designer, shows mockups, mentions Figma and user testing\"\n}\nEOF\n\n# 3. Run with both\npython3 transcribe.py meeting.wav \\\n  --lang zh --num-speakers 3 \\\n  --speakers \"Alice,Bob,Carol\" \\\n  --hotwords hotwords.txt \\\n  --speaker-context speaker-context.json\n```\n\n## Audio Preprocessing\n\nFunASR works best with 16kHz mono audio. **FLAC is recommended** over WAV — lossless\nquality at ~50% the file size, and FunASR reads it natively via soundfile.\n\n```bash\n# Recommended: FLAC (lossless, compact)\nffmpeg -i recording.m4a -ar 16000 -ac 1 -sample_fmt s16 meeting.flac\n\n# Alternative: WAV (lossless, larger)\nffmpeg -i recording.m4a -ar 16000 -ac 1 meeting.wav\n```\n\n**Important:** Use `-sample_fmt s16` when converting to FLAC — without it, ffmpeg\nmay output 24-bit samples (s32/24bit) which doubles the file size with no ASR benefit.\n\n### Format comparison (4h14m meeting)\n\n| Format | Size | Quality | FunASR support |\n|--------|------|---------|---------------|\n| **FLAC (16kHz mono s16)** | **219MB** | Lossless | Native (soundfile) |\n| WAV (16kHz mono) | 465MB | Lossless | Native (soundfile) |\n| Opus (32kbps) | 54MB | Lossy | Native (soundfile) |\n| M4A/AAC (original 48kHz) | 173MB | Source | Via librosa |\n| M4A/AAC (16kHz 32kbps) | 60MB | Lossy | Via librosa |\n\nFunASR accepts all common audio formats. FLAC offers the best trade-off: lossless\nquality, reasonable size, and native reader support without librosa fallback.\n\nFor long recordings, do NOT split the audio — FunASR handles arbitrarily long files\nand splitting breaks speaker consistency across segments.\n\n## Resume / Checkpoint Support\n\nThe pipeline supports resuming interrupted runs:\n- **Phase 1 output**: `<stem>_raw_transcript.json` — use `--skip-transcribe` to skip ASR\n- **Phase 3 cache**: `<stem>_llm_cache/chunk_NNN.txt` — already-cleaned chunks are reused\n  (kept by default; add `--clean-cache` to delete after completion)\n\n## Model Caching\n\nFunASR models (~3 GB for the `zh` preset) are downloaded from ModelScope on first run\nand cached in `~/.cache/modelscope/hub/`. On ephemeral instances (EC2, cloud VMs), the\ncache is lost when the instance is replaced, requiring a ~2 minute re-download.\n\nTo persist the cache on durable storage (e.g., an EBS data volume):\n\n```bash\n# Via CLI flag (recommended)\npython3 transcribe.py meeting.flac --model-cache-dir /data/modelscope-cache ...\n\n# Via environment variable\nMODELSCOPE_CACHE=/data/modelscope-cache python3 transcribe.py meeting.flac ...\n```\n\nThe `systemd-run` examples below include `-E MODELSCOPE_CACHE=...` for this reason.\n\n## Speaker Context JSON Format\n\nThe `--speaker-context` file helps the LLM identify speakers and fix ASR errors:\n\n```json\n{\n  \"Alice\": \"Discussed Q1 revenue targets, mentioned Chicago office relocation\",\n  \"Bob\": \"Presented the new CI/CD pipeline, uses Terraform and ArgoCD\",\n  \"Carol\": \"HR updates, mentioned hiring freeze and new PTO policy\"\n}\n```\n\nThe context is injected into the LLM system prompt for each cleanup chunk.\n\n## Running on CPU-only / Low-Memory Machines\n\nLong recordings (2+ hours) on resource-constrained machines (CPU-only, ≤8 GB RAM)\nface two common failure modes. Both are silent — the process is killed mid-run\nwith no output files saved.\n\n### Problem 1: Process killed by execution timeout\n\nAI coding agents (Claude Code, OpenClaw, Cursor, etc.) impose execution timeouts\non shell commands — typically 2–10 minutes. On a 4-hour recording, CPU transcription\ntakes 1.5–2 hours, well past any agent timeout. The process is silently killed.\n\n**Fix — detach the ASR phase from the agent's process supervision.** Use `--skip-llm`\nfor the detached run because Phase 1 (ASR) is the CPU-intensive bottleneck; Phase 3\n(LLM cleanup) is network-bound and fast — run it afterward via `--skip-transcribe`\nunder the normal agent session.\n\nOption A: `systemd-run` (preferred on systemd hosts):\n\n```bash\nsystemd-run --user --unit=transcribe-job \\\n  --working-directory=/tmp \\\n  -E MODELSCOPE_CACHE=/data/modelscope-cache \\\n  bash -c 'source /path/to/.venv/bin/activate && \\\n    python3 /path/to/transcribe.py /tmp/meeting.flac \\\n    --lang zh --num-speakers 9 --skip-llm > /tmp/transcribe.log 2>&1'\n\n# Monitor progress\nsystemctl --user status transcribe-job.service\ntail -f /tmp/transcribe.log\n\n# Check result\nls -lh /tmp/*-transcript.md /tmp/*_raw_transcript.json\n```\n\n`systemd-run` creates a transient systemd service fully independent of the agent\nsession — it survives session resets, context pruning, and exec timeouts.\n\n> **Warning:** `systemd-run` creates an isolated mount namespace. FUSE mounts\n> (rclone, sshfs, Google Drive, etc.) from the parent session are NOT visible\n> to the transient service. Copy all dependency files (audio, hotwords,\n> speaker-context, reference documents) to a local path (e.g., `/tmp`) before\n> launching. Always use `--working-directory=/tmp` or another real filesystem path.\n\n> **Note:** `systemd-run --user` requires a user-level systemd instance. It may not\n> work in Docker containers or cloud VMs without `loginctl enable-linger`.\n\nOption B: `nohup` (works everywhere):\n\n```bash\nnohup bash -c 'source .venv/bin/activate && python3 transcribe.py meeting.flac \\\n  --lang zh --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\necho $!  # Save PID for monitoring\ntail -f transcribe.log\n```\n\n### Problem 2: OOM kill on machines with ≤8 GB RAM\n\nThe `zh` preset loads 4 model components simultaneously (SeACo-Paraformer + VAD +\nPunctuation + CAM++ speaker). Peak RSS can exceed 7 GB on a 4-hour recording. On\nmachines without swap, the OOM killer terminates the process silently.\n\n**Fix A — add swap before running** (requires root):\n\n```bash\nsudo fallocate -l 4G /swapfile\nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n\n# After transcription, optionally remove swap\nsudo swapoff /swapfile && sudo rm /swapfile\n```\n\n**Fix B — use `zh-basic` instead of `zh`:**\n\n`zh-basic` (Paraformer-large) loads one fewer model component than `zh`\n(SeACo-Paraformer), reducing peak RSS by ~1–1.5 GB. Accuracy is slightly lower\n(no hotword biasing) but sufficient for most meetings:\n\n```bash\npython3 transcribe.py meeting.flac --lang zh-basic --num-speakers 9 --skip-llm\n```\n\n**Combining both fixes** (swap + `zh-basic`) reliably handles 4+ hour recordings on\nmachines with as little as 8 GB RAM + 4 GB swap.\n\n### Recommended CPU workflow for long recordings\n\n```bash\n# 1. Add swap if RAM ≤ 8 GB\nsudo fallocate -l 4G /swapfile && sudo chmod 600 /swapfile \\\n  && sudo mkswap /swapfile && sudo swapon /swapfile\n\n# 2. Launch transcription detached from agent timeout\nnohup bash -c 'source .venv/bin/activate && python3 transcribe.py meeting.flac \\\n  --lang zh-basic --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\n# 3. Monitor\ntail -f transcribe.log\n\n# 4. When done, resume with LLM cleanup (network-bound, runs fine under agent)\npython3 transcribe.py meeting.flac --skip-transcribe\n```\n\n## Podcast Transcription\n\nThe pipeline handles podcasts and interviews with the same engine, but the workflow\ndiffers from meetings:\n\n### Key differences from meetings\n\n| Aspect | Meeting | Podcast / Interview |\n|--------|---------|---------------------|\n| Speakers | 3–15+, often unknown | 2–3, usually known (host + guests) |\n| Language | Usually single | May mix languages (bilingual hosts) |\n| Hotwords | Participant names + terms | Show name, guest name, topic terms |\n| Speaker context | Role-based keywords | Host asks questions, guest answers |\n| Diarization | Critical | Easier (fewer, distinct voices) |\n\n### Recommended settings\n\n```bash\n# English podcast (2 speakers, host + guest)\npython3 transcribe.py episode.flac --lang en --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n\n# Bilingual podcast (auto-detect language switches)\npython3 transcribe.py episode.flac --lang auto --num-speakers 2 \\\n  --speakers \"Alice,Bob\"\n\n# Chinese podcast with topic hotwords\npython3 transcribe.py episode.flac --lang zh --num-speakers 3 \\\n  --speakers \"主持人,嘉宾A,嘉宾B\" \\\n  --hotwords \"播客名 嘉宾全名 讨论主题关键词\"\n\n# Multi-language podcast (e.g., Spanish + English)\npython3 transcribe.py episode.flac --lang whisper --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n```\n\n### Tips for podcast transcription\n\n1. **Always provide `--num-speakers`** — podcasts have a known, fixed speaker count;\n   this dramatically improves diarization accuracy with only 2–3 voices\n2. **Always provide `--speakers`** — host/guest names are known upfront\n3. **`--lang auto`** works for bilingual transcript-only output (no speaker labels) —\n   SenseVoiceSmall handles intra-utterance language switching (zh/en/ja/ko/yue) but\n   does not output timestamps, so **speaker diarization is not supported**.\n   Use `--lang zh` for Chinese podcasts that need speaker identification.\n4. **`--lang whisper`** for any other language or heavy code-switching (also lacks\n   timestamp support for diarization — transcript-only)\n5. **Hotwords** — for Chinese podcasts, include the show name and guest's full name;\n   for English podcasts, hotwords are usually unnecessary\n6. **`--speaker-context`** — describe the host/guest dynamic:\n   ```json\n   {\n     \"Alice\": \"Host, asks questions, introduces topics, wraps up segments\",\n     \"Bob\": \"Guest, expert on topic X, shares personal anecdotes\"\n   }\n   ```\n7. **Audio quality** — podcasts are typically studio-recorded with better SNR than\n   meetings; diarization accuracy is correspondingly higher\n8. **Audio source matters** — mobile app downloads may be truncated (trial/preview\n   versions). Download from the web interface for complete files. The script's\n   Phase 0 duration validation catches conversion truncation but cannot detect\n   a source file that is already incomplete.\n\n## `--lang mimo` — Xiaomi MiMo-V2.5-ASR (local, GPU)\n\nMiMo is an 8B-parameter LLM-based ASR model from Xiaomi. It outputs plain text\nwith no per-sentence timestamps and no speaker labels. The `audio-transcribe`\nskill wraps it in a VAD + speaker-clustering sandwich so output format matches\nthe FunASR presets:\n\n```\nPhase 1a  FSMN VAD           → [(start_ms, end_ms), ...]\nPhase 1b  MiMo asr_sft()     → text per VAD segment\nPhase 1c  CAM++ + KMeans     → speaker ID per VAD segment\n```\n\nFiles: `scripts/mimo_asr.py` (orchestrator), `scripts/setup_mimo.sh` (installer).\n\n### Expected RTF\n\nOn a single A100 (40 GB), 4h audio → ~24 min wall clock for Phase 1 (RTF ≈\n0.1). Compare to `--lang zh` on the same GPU at RTF ≈ 0.02–0.05. Trade speed\nfor reported accuracy gains on dialects, code-switching, and lyrics.\n\n### Resume (`--resume-mimo`)\n\nSegment-level failures (OOM, CUDA error) retry 3× with\n[0.5s, 2s, 5s] backoff after `gc.collect()` + `torch.cuda.empty_cache()`. If all\nretries fail, a `*_mimo_partial.json` file captures VAD segments, completed\ntranscriptions, and the failed index. `--resume-mimo` picks up from the failed\nsegment, verifying audio SHA256 + `--mimo-audio-tag` match before continuing.\n\nFile v1.7.1:skill-card.md\n\n## Description:\n\nAudio Transcribe helps an agent transcribe meeting, podcast, and interview recordings into structured transcripts with ASR, speaker diarization, optional speaker verification, and optional LLM cleanup.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zxkane](https://clawhub.ai/user/zxkane)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, operators, and external users can use this skill to convert audio recordings into Markdown transcripts with speaker labels, timestamps, and optional cleanup for meeting notes or podcast/interview review.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Sensitive audio, transcript excerpts, speaker context, or reference material may be sent to external LLM providers when cleanup is enabled.\n\nMitigation: Omit --model or use --skip-llm to keep cleanup local, and only provide trusted reference and speaker-context files when LLM cleanup is needed.\n\nRisk: Speaker-gender inference is enabled by default and may create unnecessary demographic processing.\n\nMitigation: Use --no-detect-gender when demographic inference is not required.\n\nRisk: OpenAI-compatible endpoints and setup commands can change the operational trust boundary.\n\nMitigation: Review any OPENAI_BASE_URL value and run setup commands in an isolated virtual environment after reviewing privileged and dependency behavior.\n\n## Reference(s):\n\n- [Project homepage](https://github.com/zxkane/audio-transcriber)\n- [Pipeline Details](references/pipeline-details.md)\n- [Xiaomi MiMo-V2.5-ASR Model Card](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR)\n- [ClawHub Skill Page](https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, JSON, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands; generated transcript output may include Markdown and raw JSON files.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Transcription output can include speaker labels, timestamps, hotword-biased ASR results, and optional LLM-cleaned text.]\n\n## Skill Version(s):\n\n1.7.1 (source: SKILL.md frontmatter and server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.7.0: 13 files, 85490 bytes\n\nFiles: references/pipeline-details.md (20640b), scripts/llm_utils.py (7469b), scripts/mimo_asr.py (19141b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (5882b), scripts/setup_mimo.sh (6132b), scripts/speaker_gender.py (11457b), scripts/test_mimo_asr.py (24599b), scripts/test_speaker_verification.py (92490b), scripts/transcribe_funasr.py (68763b), scripts/verify_speakers.py (18210b), SKILL.md (15390b), _meta.json (150b)\n\nFile v1.7.0:SKILL.md\n\n---\nname: funasr-transcribe\nversion: 1.7.0\ndescription: >\n  This skill should be used when the user explicitly asks to \"transcribe a meeting\",\n  \"transcribe audio\", \"transcribe a meeting recording\",\n  \"convert audio to text\", \"generate meeting minutes from audio\",\n  \"do speech-to-text\", \"transcribe with speaker diarization\",\n  \"identify speakers in audio\", \"transcribe Chinese audio\",\n  \"transcribe English audio\", \"transcribe Japanese audio\",\n  \"multi-speaker transcription\", \"transcribe a podcast\",\n  \"transcribe podcast episode\", \"transcribe an interview\",\n  \"convert podcast to text\", \"podcast to transcript\",\n  or mentions FunASR, Paraformer, SenseVoice, Whisper, meeting\n  transcription, podcast transcription, or speaker diarization.\n  Supports multi-speaker meeting and podcast transcription in Chinese,\n  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper)\n  with automatic speaker diarization and hotword biasing.\n  Works on both GPU and CPU.\nmetadata:\n  openclaw:\n    requires:\n      bins: [\"python3\", \"ffmpeg\"]\n    env_vars:\n      - name: AWS_REGION\n        required: false\n        description: \"AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed.\"\n      - name: ANTHROPIC_API_KEY\n        required: false\n        description: \"API key for Anthropic Claude LLM cleanup\"\n      - name: OPENAI_API_KEY\n        required: false\n        description: \"API key for OpenAI-compatible LLM cleanup\"\n      - name: OPENAI_BASE_URL\n        required: false\n        description: \"Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)\"\n    emoji: \"🎙️\"\n    homepage: \"https://github.com/zxkane/audio-transcriber-funasr\"\n---\n\n# FunASR Meeting & Podcast Transcription\n\nTranscribe multi-speaker audio into structured Markdown with automatic\nspeaker diarization, hotword biasing, and optional LLM cleanup.\n\nAll scripts run directly from the plugin directory — no copying needed.\nDefine this shorthand at the start of every session:\n\n```bash\nSCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/funasr-transcribe/scripts\n```\n\n## Supported Languages\n\n| `--lang` | Model | Languages | Hotword |\n|----------|-------|-----------|---------|\n| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |\n| `zh-basic` | Paraformer-large | Chinese | No |\n| `en` | Paraformer-en | English | No |\n| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |\n| `whisper` | Whisper-large-v3-turbo | 99 languages | No |\n| `mimo` | MiMo-V2.5-ASR (local 8B, GPU-only) | zh/en/code-switch/dialects | No |\n\nAll presets include **speaker diarization** (CAM++) and **VAD** (FSMN).\n`mimo` reuses the FSMN VAD + CAM++ stack around MiMo's text output.\n\n> **Diarization caveat:** `auto` and `whisper` do not output per-sentence timestamps,\n> so speaker diarization does not work with these presets. Use `zh`, `zh-basic`,\n> `en`, or `mimo` when speaker identification is needed (e.g., podcasts, meetings).\n\n## Workflow\n\nBefore starting transcription, **always ask the user**:\n\n1. **Audio file** — path to the recording (required)\n2. **Type** — meeting, podcast, or interview? (affects defaults)\n3. **Language** — what language is spoken? (default: Chinese)\n4. **Number of speakers** — how many participants? (improves diarization)\n5. **Speaker names** — for podcasts: host + guest names; for meetings: attendee list\n6. **Supporting files** — ask:\n   > \"Do you have any of the following to improve accuracy?\"\n   > - **Attendee / guest list** — for hotwords and speaker mapping\n   > - **Meeting agenda or episode topic** — for hotwords (terms, names)\n   > - **Reference documents** (show notes, prior notes) — for speaker identification and ASR correction\n\n**Adapt defaults by recording type:**\n- **Meeting**: default `--lang zh`, ask about supporting files\n- **Podcast / interview**: default `--lang zh`, `--num-speakers 2`, always ask for\n  host + guest names, suggest `--speaker-context` for roles\n  (do NOT use `--lang auto` — it lacks timestamps for speaker diarization)\n\n> **⚠️ `--speakers` must use the speaker's real name, not a podcast alias.**\n> The value passed to `--speakers` is used verbatim as the speaker label in the\n> output transcript. Always derive it from the host/guest's actual name (e.g.\n> from a shownotes \"Host:\" field), not from the podcast feed name or title.\n>\n> Example: if shownotes lists \"Host: 张三（张三的播客）\", pass `--speakers '张三'`\n> — not the alias \"张三的播客\". Add both the real name and the alias to\n> `hotwords.txt` so ASR can recognise both forms.\n>\n> When both `--speakers` and `--reference` are supplied, the script detects\n> this mistake at startup and prints an `ACTION REQUIRED` block naming the\n> suggested real name. **If you see that block, stop the run and re-invoke\n> with the corrected `--speakers` value before Phase 3** — the warning does\n> not abort the pipeline.\n\nIf the user provides supporting materials:\n- Extract participant names and key terms → create `hotwords.txt` (include both real name and alias)\n- Extract per-person context → create `speaker-context.json`\n- Pass original reference document with `--reference`\n- Use all three together for best results\n\n## Quick Start\n\n### 1. Environment Setup\n\n```bash\nAUTO_YES=1 bash $SCRIPTS/setup_env.sh\n# Or force CPU:  AUTO_YES=1 bash $SCRIPTS/setup_env.sh cpu\n```\n\nThe setup script patches FunASR's spectral clustering for O(N²·k) performance.\nWithout this, recordings over ~1 hour hang for hours during speaker clustering.\n\n### 2. Run Transcription\n\nOutput files are written to the current working directory.\n\n**LLM cleanup (Phase 3) is opt-in.** By default, transcription runs locally\nwithout contacting any external service. To enable LLM-powered ASR correction\nand speaker name refinement, pass `--model <model-id>`. Use LLM cleanup when:\n- The raw transcript has many ASR errors (names, technical terms)\n- You need polished, publication-ready output\n- Speaker names need to be refined from context\n\n> **⚠️ Data Privacy:** When LLM cleanup is enabled via `--model`, transcript\n> excerpts are sent to external LLM providers (AWS Bedrock, Anthropic, or\n> OpenAI depending on the model ID). Use `--skip-llm` or omit `--model` to\n> keep all data local. For Bedrock, boto3 uses the standard AWS credential\n> chain (IAM role, SSO, `~/.aws/credentials`, env vars).\n\n```bash\n# Chinese meeting with hotwords (local-only, no LLM)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt\n\n# English meeting with speaker names\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang en --speakers \"Alice,Bob,Carol,Dave\"\n\n# Auto-detect language (zh/en/ja/ko/yue)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang auto --num-speakers 6\n\n# Whisper for any language\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang whisper --num-speakers 4\n\n# Enable LLM cleanup for polished output (requires --model)\n# Bedrock (uses AWS credential chain: IAM role, SSO, ~/.aws/credentials)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt \\\n    --provider bedrock --model us.anthropic.claude-sonnet-4-6\n\n# Bedrock \"global\" cross-region profile (recent AWS deployments)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --provider bedrock --model global.anthropic.claude-sonnet-4-6\n\n# Bedrock via litellm-style wrapper (supported; prefix is stripped for boto3)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --provider bedrock --model amazon-bedrock/global.anthropic.claude-sonnet-4-6\n\n# Anthropic API (requires ANTHROPIC_API_KEY env var)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --provider anthropic --model claude-sonnet-4-6\n\n# OpenAI-compatible API (requires OPENAI_API_KEY env var)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --provider openai --model gpt-4o\n\n# Full pipeline with all supporting files + LLM (best quality)\npython3 $SCRIPTS/transcribe_funasr.py episode.m4a \\\n    --lang zh --num-speakers 2 \\\n    --hotwords hotwords.txt \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json \\\n    --reference show-notes.md \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Resume interrupted LLM cleanup\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --skip-transcribe --model us.anthropic.claude-sonnet-4-6\n```\n\n### 3. Verify Speaker Labels\n\nIf the transcript has swapped speaker labels (common with podcasts),\nthe verification script can detect and fix mismatches using LLM analysis:\n\n```bash\n# Dry-run: check if host/guest are swapped\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json\n\n# Apply the fix\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json --fix\n\n# Multi-speaker meeting: full reassignment\npython3 $SCRIPTS/verify_speakers.py meeting_raw_transcript.json \\\n    --speakers \"Alice,Bob,Carol,Dave\" \\\n    --speaker-context speaker-context.json --fix\n\n# Then regenerate the markdown with corrected labels\npython3 $SCRIPTS/transcribe_funasr.py original.m4a \\\n    --skip-transcribe --clean-cache\n```\n\nThe script analyzes the first 5 minutes (configurable with `--minutes`)\nand auto-detects podcast (2 speakers, swap detection) vs meeting\n(N speakers, full reassignment).\n\n## Audio Preprocessing\n\nThe script automatically converts input audio to 16kHz mono FLAC and\nvalidates that no audio is lost (detects silent truncation).\n\n| Format | 4h14m meeting | Quality | Recommendation |\n|--------|--------------|---------|----------------|\n| **FLAC** | **219MB** | Lossless | **Default, safest** |\n| Opus | 55MB | Lossy | Risk of truncation on long files |\n| WAV | 465MB | Lossless | Works but larger |\n| Original M4A | 173MB | Source | Also works directly |\n\n**Do NOT split long recordings** — splitting breaks speaker ID consistency.\n\n## MiMo-V2.5-ASR (optional, GPU-only)\n\n`--lang mimo` runs Xiaomi's\n[MiMo-V2.5-ASR](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR) locally on a\nCUDA GPU. Use it when:\n- You want to evaluate MiMo against Paraformer on Chinese audio.\n- The recording has heavy code-switching, dialects (Wu, Cantonese, Hokkien,\n  Sichuanese), lyrics, or rare proper nouns that other presets mis-transcribe.\n\n**Requirements:**\n- CUDA ≥12.0 and **≥20 GB VRAM** (16 GB cards OOM during inference).\n- Python 3.12 (enforced by `setup_env.sh`).\n- ~20 GB weight download (one-time) and `flash-attn==2.7.4.post1` compile\n  (needs `nvcc` from the CUDA toolkit, takes 10–30 min).\n\n**Install (opt-in):**\n\n```bash\n# One-time: install MiMo on top of the standard environment\nAUTO_YES=1 INSTALL_MIMO=1 \\\n    MIMO_WEIGHTS_PATH=/mnt/models/hf \\\n    bash $SCRIPTS/setup_env.sh\n```\n\n**Run:**\n\n```bash\npython3 $SCRIPTS/transcribe_funasr.py podcast.m4a \\\n    --lang mimo --num-speakers 2 \\\n    --mimo-weights-path /mnt/models/hf\n```\n\n**Resume after failure:**\n\n```bash\npython3 $SCRIPTS/transcribe_funasr.py podcast.m4a \\\n    --lang mimo --resume-mimo --mimo-weights-path /mnt/models/hf\n```\n\n**Limitations:**\n- No hotword biasing (MiMo has no API for it — `--hotwords` is ignored).\n- No CPU fallback.\n- Inference is slower than Paraformer on the same GPU (8B model vs ~0.3B);\n  expect RTF around 0.1–0.2 on an A100.\n\n## Key Flags\n\n| Flag | Purpose |\n|------|---------|\n| `--lang` | `zh` (default), `zh-basic`, `en`, `auto`, `whisper` |\n| `--hotwords` | Hotword file or string — biases ASR (zh only) |\n| `--reference F` | Reference file for LLM ASR correction |\n| `--num-speakers N` | Expected speaker count (improves diarization) |\n| `--speakers \"A,B,C\"` | Assign real names by first-appearance order |\n| `--speaker-context F` | JSON with per-speaker roles for LLM |\n| `--no-detect-gender` | Disable automatic speaker gender detection (CAM++ gender classifier) |\n| `--speaker-genders \"A:female,B:male\"` | Override per-speaker gender (also accepts positional `female,male`) |\n| `--audio-format` | `flac` (default), `opus`, `wav` |\n| `--device cpu` | Force CPU mode |\n| `--batch-size N` | Adjust for memory (60 for CPU, 100 if GPU OOM) |\n| `--phase1-only` | Exit after Phase 1 (VAD + ASR + diarization), skip Phase 2 + 3 |\n| `--json-out PATH` | Write raw transcript JSON to explicit path (overrides default naming) |\n| `--skip-transcribe` | Resume from saved `*_raw_transcript.json` |\n| `--skip-llm` | Skip LLM cleanup (default when `--model` is omitted) |\n| `--model ID` | Enable LLM cleanup with this model (auto-detects Bedrock/Anthropic/OpenAI) |\n| `--title \"...\"` | Output document title |\n| `--clean-cache` | Delete LLM chunk cache after completion |\n| `--output PATH` | Custom output file path |\n| `--model-cache-dir` | ModelScope model cache directory (~3GB, default: `~/.cache/modelscope/`) |\n| `--mimo-audio-tag` | MiMo language hint: `<chinese>` (default), `<english>`, `<auto>` |\n| `--mimo-batch N` | Concurrent VAD segments per MiMo call (default 1; H100/80GB can go higher) |\n| `--mimo-weights-path DIR` | Cache dir for MiMo weights (default: `$HF_HOME` → `~/.cache/huggingface`) |\n| `--resume-mimo` | Resume MiMo Phase 1 from `*_mimo_partial.json` after a mid-run failure |\n\n## Outputs\n\n- `<stem>-transcript.md` — Final Markdown with speaker labels and timestamps\n- `<stem>_raw_transcript.json` — Raw Phase 1 output (for resume/analysis)\n\n## Speaker Diarization Tips\n\nFunASR's CAM++ may merge acoustically similar speakers. To improve:\n\n1. **`--num-speakers N`** — Hint expected count\n2. **`--hotwords`** — Include participant names (Chinese names work best)\n3. **`--speaker-context`** — Provide per-person keywords for LLM splitting\n4. **Keyword matching** — Search `*_raw_transcript.json` for unique phrases\n\n### Speaker gender\n\nEnabled by default: each detected speaker is classified as `male` / `female`\nvia 3D-Speaker's CAM++ gender classifier (`iic/speech_campplus_two_class_gender_16k`).\nThe result appears next to each name in the **Speaker List** table and is\ninjected into the LLM cleanup prompt so pronouns (他/她, he/she) get corrected.\n\nPrecedence when combined:\n\n1. `--speaker-genders \"Alice:female,Bob:male\"` (explicit CLI) — always wins\n2. Reference text hints like `主播（女）：韩梅梅` or `Host (male): Alice` — override auto\n3. CAM++ auto-detection — fallback\n\nDisable with `--no-detect-gender` if you don't need gender and want to save\nthe ~500 MB model download and extra inference time.\n\n## CPU-only / Low-Memory Machines\n\nLong recordings on resource-constrained machines may hit exec timeouts\nor OOM kills. See `references/pipeline-details.md` for workarounds:\n- Detach from agent timeouts with `systemd-run` or `nohup`\n- Prevent OOM via swap and/or `--lang zh-basic` (lighter model)\n\n## Additional Resources\n\n- **`references/pipeline-details.md`** — Architecture, model specs, benchmarks,\n  speaker role verification, hotword effectiveness, clustering patch\n- **`scripts/transcribe_funasr.py`** — Main transcription pipeline\n- **`scripts/verify_speakers.py`** — Speaker label verification & fix\n- **`scripts/llm_utils.py`** — Shared LLM infrastructure (Bedrock/Anthropic/OpenAI)\n- **`scripts/setup_env.sh`** — Environment setup (venv + deps + patch)\n\nFile v1.7.0:_meta.json\n\n{\n  \"ownerId\": \"kn72agp4n0v3y1gk4qds89r0wn82msn3\",\n  \"slug\": \"zxkane-audio-transcriber-funasr\",\n  \"version\": \"1.7.0\",\n  \"publishedAt\": 1777596125731\n}\n\nFile v1.7.0:references/pipeline-details.md\n\n# FunASR Meeting Transcription Pipeline — Technical Details\n\n## Architecture\n\n```\nAudio File (.m4a/.mp3/.wav)\n  │\n  ├─ [ffmpeg] ──► 16kHz mono WAV\n  │\n  ├─ [Phase 1: FunASR] ──► raw_transcript.json\n  │   ├─ FSMN-VAD: segment speech vs silence\n  │   ├─ ASR model (language-dependent, see below)\n  │   ├─ (Optional) Hotword biasing (SeACo-Paraformer only)\n  │   ├─ Punctuation restoration (model-dependent)\n  │   └─ CAM++: speaker embeddings → spectral clustering\n  │\n  ├─ [Phase 2: Post-process]\n  │   ├─ Merge consecutive same-speaker utterances (<2s gap)\n  │   ├─ Map speaker IDs to names (if provided)\n  │   └─ Auto-verify via self-introduction detection\n  │\n  └─ [Phase 3: LLM cleanup] ──► transcript.md\n      ├─ LLM speaker role verification (if --speaker-context provided)\n      ├─ Remove fillers (um, uh, 嗯, 啊, etc.)\n      ├─ Fix ASR errors (homophones, context-based)\n      ├─ Polish grammar while preserving meaning\n      └─ (Optional) Identify merged speakers via context\n```\n\n## Language Presets & Models\n\n### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |\n| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |\n| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |\n| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |\n\nSeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)\nto bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.\n\n### `--lang zh-basic` (Chinese, no hotword)\n\nSame as `zh` but uses the base Paraformer-large without hotword support.\nUse when hotword biasing is unnecessary or causing issues with English terms.\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |\n\n### `--lang en` (English)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |\n\n### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |\n\nSenseVoiceSmall includes built-in punctuation and supports emotion detection.\n\n### `--lang whisper` (Multilingual, 99 languages)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |\n\nAll presets share the same VAD (`fsmn-vad`) and speaker diarization (`cam++`) models.\nModels are auto-downloaded from ModelScope on first run.\n\n## Hotword Biasing\n\nSeACo-Paraformer (`--lang zh`) supports hotword customization to improve recognition\nof specific terms — particularly useful for participant names, project names, and\ndomain-specific jargon in meetings.\n\n### How to provide hotwords\n\n```bash\n# Space-separated string\npython3 transcribe_funasr.py meeting.wav --lang zh --hotwords \"张三 李四 ClawCon Rebase\"\n\n# Text file (one word per line)\npython3 transcribe_funasr.py meeting.wav --lang zh --hotwords hotwords.txt\n```\n\n### What to include in hotwords\n\nFor meeting transcription, a good hotwords file includes:\n- **Participant names** (full names in the meeting's language)\n- **Project / product names** (internal codenames, product brands)\n- **Domain-specific Chinese terms** (technical jargon, acronyms in Chinese)\n- **Organization names** (company, team, department names)\n\n### Effectiveness — empirical results\n\nTested on a 4h14m, 9-speaker Chinese meeting with 27 hotwords:\n\n| Term | Without hotwords | With hotwords | Change |\n|------|-----------------|---------------|--------|\n| 龙虾 (lobechat) | 28 | 42 | **+50%** |\n| 高琦 (person name) | 0 | 7 | **0 → 7** |\n| 搬瓦工 (BandwagonHost) | 0 | 1 | **0 → 1** |\n| 谢锐 (person name) | 0 | 1 | improved |\n| 鲲鹏 (org name) | 6 | 7 | slight improvement |\n| Rebase (English) | 5 | 0 | **regression** |\n| Tailwind (English) | 3 | 1 | **regression** |\n\n**Key findings:**\n1. **Chinese terms benefit significantly** — names, brands, and Chinese jargon\n   see clear improvement (龙虾 +50%, 高琦 from zero)\n2. **English terms may regress** — SeACo's hotword biasing operates on Chinese\n   token vocabulary; English loanwords can be disrupted\n3. **Person names have limited uplift** — meeting participants rarely say full\n   names aloud; hotwords help only when names do appear in speech\n\n**Recommendation:** Include Chinese terms and names in hotwords. For English\ntechnical terms, rely on Phase 3 LLM cleanup rather than hotword biasing.\nIf English term accuracy is critical and hotword biasing causes regressions,\nuse `--lang zh-basic` instead.\n\n## Performance Benchmarks\n\nTested on a 4h14m, 9-speaker Chinese meeting recording (GPU: L40S 46GB):\n\n| Metric | Paraformer-large | SeACo + Hotword | Notes |\n|--------|-----------------|-----------------|-------|\n| Model load | 14s | 14s (422s first run, downloading 944MB model) | Cached after first run |\n| Transcription | 169s | 168s | Virtually identical |\n| Raw sentences | 6672 | 6725 | Comparable |\n| Merged segments | 1695 | 1724 | Comparable |\n| Speakers detected | 7 (of 9) | 7 (of 9) | Same diarization result |\n\n| Metric | GPU (L40S 46GB) | CPU (estimated) |\n|--------|-----------------|-----------------|\n| Model load | 14s | ~30s |\n| Transcription | 169s | ~30-60 min |\n| Speaker clustering | ~10s (patched) | ~2-5 min (patched) |\n| LLM cleanup (17 chunks) | ~35 min | ~35 min (network-bound) |\n| Total | ~38 min | ~70-100 min |\n\n**Without the clustering patch**, the original `scipy.linalg.eigh()` on the full Laplacian\nmatrix was O(N^3) and took **10+ hours** on this recording. The patch reduces it to O(N^2*k)\nvia `scipy.sparse.linalg.eigsh()`.\n\n## Clustering Patch (Critical for Long Meetings)\n\nFunASR's `SpectralCluster.get_spec_embs()` uses `scipy.linalg.eigh(L)` which computes\nALL eigenvalues of the NxN Laplacian. For a 4-hour recording, N can be 6000+, making\nthis O(N^3) operation take hours.\n\nThe patch (`scripts/patch_clustering.py`) replaces this with:\n- `scipy.sparse.linalg.eigsh(L_sparse, k=num_speakers, which='SM')` — only computes\n  the k smallest eigenvalues needed, reducing complexity to O(N^2 * k)\n- Vectorized `p_pruning()` — replaces Python loop with numpy broadcasting\n\n**Always run the patch before processing meetings longer than ~1 hour.**\n\n## Speaker Role Verification\n\nSpeaker names are assigned by first-appearance order in the audio, which may\nswap host/guest labels (especially in podcasts). Two layers of automatic\nverification, plus a standalone post-hoc tool:\n\n1. **Phase 2 — self-introduction detection**: scans the first 5 minutes for\n   explicit self-introductions (\"我是X\", \"I'm X\") and swaps labels if mismatched\n2. **Phase 3 — LLM role verification**: when `--speaker-context` is provided,\n   the LLM analyzes the first chunk (up to 15 minutes) before cleanup begins.\n   For 2 speakers: binary CORRECT/SWAP detection. For 3+ speakers: full\n   JSON-based reassignment matching each label to the correct person.\n3. **Post-hoc — `verify_speakers.py`**: standalone script that verifies any\n   existing `*_raw_transcript.json`. Same two modes (2-speaker swap, N-speaker\n   reassignment) with dry-run support. See SKILL.md § Verify Speaker Labels.\n\nFor podcasts, always provide `--speaker-context` describing host/guest roles.\n\n## Speaker Diarization Limitations\n\nFunASR's CAM++ speaker diarization may merge acoustically similar speakers into one ID.\nIn the tested 9-person meeting, only 7 unique IDs were detected (two pairs merged).\n\nWorkarounds:\n1. **Provide `--num-speakers N`** to hint expected count (uses `preset_spk_num`)\n2. **Post-hoc keyword matching**: use reference documents (meeting agendas, attendee notes)\n   to identify which speaker ID maps to which person\n3. **LLM-assisted splitting**: provide `--speaker-context` with per-person keywords;\n   the LLM can then split merged speakers when context is clear (~73% success rate)\n\n## Supporting Files for Better Results\n\nPrepare these files before transcription for best results:\n\n| File | Used in | Purpose |\n|------|---------|---------|\n| `hotwords.txt` | Phase 1 (`--hotwords`) | Bias ASR toward names and terms |\n| `speaker-context.json` | Phase 3 (`--speaker-context`) | Help LLM identify and split speakers |\n| Meeting agenda | Manual reference | Identify meeting phases for post-analysis |\n| Attendee list | Build hotwords + speaker names | Map speaker IDs to real names |\n\n### Example: preparing supporting files from a meeting invite\n\n```bash\n# 1. Create hotwords.txt from attendee list and agenda\ncat > hotwords.txt << 'EOF'\nAlice\nBob\nCarol\nProjectAlpha\nSprint Review\nQ2 OKR\nEOF\n\n# 2. Create speaker-context.json from attendee roles\ncat > speaker-context.json << 'EOF'\n{\n  \"Alice\": \"Engineering manager, discusses sprint velocity and tech debt\",\n  \"Bob\": \"Product manager, presents roadmap and customer feedback\",\n  \"Carol\": \"Designer, shows mockups, mentions Figma and user testing\"\n}\nEOF\n\n# 3. Run with both\npython3 transcribe_funasr.py meeting.wav \\\n  --lang zh --num-speakers 3 \\\n  --speakers \"Alice,Bob,Carol\" \\\n  --hotwords hotwords.txt \\\n  --speaker-context speaker-context.json\n```\n\n## Audio Preprocessing\n\nFunASR works best with 16kHz mono audio. **FLAC is recommended** over WAV — lossless\nquality at ~50% the file size, and FunASR reads it natively via soundfile.\n\n```bash\n# Recommended: FLAC (lossless, compact)\nffmpeg -i recording.m4a -ar 16000 -ac 1 -sample_fmt s16 meeting.flac\n\n# Alternative: WAV (lossless, larger)\nffmpeg -i recording.m4a -ar 16000 -ac 1 meeting.wav\n```\n\n**Important:** Use `-sample_fmt s16` when converting to FLAC — without it, ffmpeg\nmay output 24-bit samples (s32/24bit) which doubles the file size with no ASR benefit.\n\n### Format comparison (4h14m meeting)\n\n| Format | Size | Quality | FunASR support |\n|--------|------|---------|---------------|\n| **FLAC (16kHz mono s16)** | **219MB** | Lossless | Native (soundfile) |\n| WAV (16kHz mono) | 465MB | Lossless | Native (soundfile) |\n| Opus (32kbps) | 54MB | Lossy | Native (soundfile) |\n| M4A/AAC (original 48kHz) | 173MB | Source | Via librosa |\n| M4A/AAC (16kHz 32kbps) | 60MB | Lossy | Via librosa |\n\nFunASR accepts all common audio formats. FLAC offers the best trade-off: lossless\nquality, reasonable size, and native reader support without librosa fallback.\n\nFor long recordings, do NOT split the audio — FunASR handles arbitrarily long files\nand splitting breaks speaker consistency across segments.\n\n## Resume / Checkpoint Support\n\nThe pipeline supports resuming interrupted runs:\n- **Phase 1 output**: `<stem>_raw_transcript.json` — use `--skip-transcribe` to skip ASR\n- **Phase 3 cache**: `<stem>_llm_cache/chunk_NNN.txt` — already-cleaned chunks are reused\n  (kept by default; add `--clean-cache` to delete after completion)\n\n## Model Caching\n\nFunASR models (~3 GB for the `zh` preset) are downloaded from ModelScope on first run\nand cached in `~/.cache/modelscope/hub/`. On ephemeral instances (EC2, cloud VMs), the\ncache is lost when the instance is replaced, requiring a ~2 minute re-download.\n\nTo persist the cache on durable storage (e.g., an EBS data volume):\n\n```bash\n# Via CLI flag (recommended)\npython3 transcribe_funasr.py meeting.flac --model-cache-dir /data/modelscope-cache ...\n\n# Via environment variable\nMODELSCOPE_CACHE=/data/modelscope-cache python3 transcribe_funasr.py meeting.flac ...\n```\n\nThe `systemd-run` examples below include `-E MODELSCOPE_CACHE=...` for this reason.\n\n## Speaker Context JSON Format\n\nThe `--speaker-context` file helps the LLM identify speakers and fix ASR errors:\n\n```json\n{\n  \"Alice\": \"Discussed Q1 revenue targets, mentioned Chicago office relocation\",\n  \"Bob\": \"Presented the new CI/CD pipeline, uses Terraform and ArgoCD\",\n  \"Carol\": \"HR updates, mentioned hiring freeze and new PTO policy\"\n}\n```\n\nThe context is injected into the LLM system prompt for each cleanup chunk.\n\n## Running on CPU-only / Low-Memory Machines\n\nLong recordings (2+ hours) on resource-constrained machines (CPU-only, ≤8 GB RAM)\nface two common failure modes. Both are silent — the process is killed mid-run\nwith no output files saved.\n\n### Problem 1: Process killed by execution timeout\n\nAI coding agents (Claude Code, OpenClaw, Cursor, etc.) impose execution timeouts\non shell commands — typically 2–10 minutes. On a 4-hour recording, CPU transcription\ntakes 1.5–2 hours, well past any agent timeout. The process is silently killed.\n\n**Fix — detach the ASR phase from the agent's process supervision.** Use `--skip-llm`\nfor the detached run because Phase 1 (ASR) is the CPU-intensive bottleneck; Phase 3\n(LLM cleanup) is network-bound and fast — run it afterward via `--skip-transcribe`\nunder the normal agent session.\n\nOption A: `systemd-run` (preferred on systemd hosts):\n\n```bash\nsystemd-run --user --unit=transcribe-job \\\n  --working-directory=/tmp \\\n  -E MODELSCOPE_CACHE=/data/modelscope-cache \\\n  bash -c 'source /path/to/.venv/bin/activate && \\\n    python3 /path/to/transcribe_funasr.py /tmp/meeting.flac \\\n    --lang zh --num-speakers 9 --skip-llm > /tmp/transcribe.log 2>&1'\n\n# Monitor progress\nsystemctl --user status transcribe-job.service\ntail -f /tmp/transcribe.log\n\n# Check result\nls -lh /tmp/*-transcript.md /tmp/*_raw_transcript.json\n```\n\n`systemd-run` creates a transient systemd service fully independent of the agent\nsession — it survives session resets, context pruning, and exec timeouts.\n\n> **Warning:** `systemd-run` creates an isolated mount namespace. FUSE mounts\n> (rclone, sshfs, Google Drive, etc.) from the parent session are NOT visible\n> to the transient service. Copy all dependency files (audio, hotwords,\n> speaker-context, reference documents) to a local path (e.g., `/tmp`) before\n> launching. Always use `--working-directory=/tmp` or another real filesystem path.\n\n> **Note:** `systemd-run --user` requires a user-level systemd instance. It may not\n> work in Docker containers or cloud VMs without `loginctl enable-linger`.\n\nOption B: `nohup` (works everywhere):\n\n```bash\nnohup bash -c 'source .venv/bin/activate && python3 transcribe_funasr.py meeting.flac \\\n  --lang zh --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\necho $!  # Save PID for monitoring\ntail -f transcribe.log\n```\n\n### Problem 2: OOM kill on machines with ≤8 GB RAM\n\nThe `zh` preset loads 4 model components simultaneously (SeACo-Paraformer + VAD +\nPunctuation + CAM++ speaker). Peak RSS can exceed 7 GB on a 4-hour recording. On\nmachines without swap, the OOM killer terminates the process silently.\n\n**Fix A — add swap before running** (requires root):\n\n```bash\nsudo fallocate -l 4G /swapfile\nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n\n# After transcription, optionally remove swap\nsudo swapoff /swapfile && sudo rm /swapfile\n```\n\n**Fix B — use `zh-basic` instead of `zh`:**\n\n`zh-basic` (Paraformer-large) loads one fewer model component than `zh`\n(SeACo-Paraformer), reducing peak RSS by ~1–1.5 GB. Accuracy is slightly lower\n(no hotword biasing) but sufficient for most meetings:\n\n```bash\npython3 transcribe_funasr.py meeting.flac --lang zh-basic --num-speakers 9 --skip-llm\n```\n\n**Combining both fixes** (swap + `zh-basic`) reliably handles 4+ hour recordings on\nmachines with as little as 8 GB RAM + 4 GB swap.\n\n### Recommended CPU workflow for long recordings\n\n```bash\n# 1. Add swap if RAM ≤ 8 GB\nsudo fallocate -l 4G /swapfile && sudo chmod 600 /swapfile \\\n  && sudo mkswap /swapfile && sudo swapon /swapfile\n\n# 2. Launch transcription detached from agent timeout\nnohup bash -c 'source .venv/bin/activate && python3 transcribe_funasr.py meeting.flac \\\n  --lang zh-basic --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\n# 3. Monitor\ntail -f transcribe.log\n\n# 4. When done, resume with LLM cleanup (network-bound, runs fine under agent)\npython3 transcribe_funasr.py meeting.flac --skip-transcribe\n```\n\n## Podcast Transcription\n\nThe pipeline handles podcasts and interviews with the same engine, but the workflow\ndiffers from meetings:\n\n### Key differences from meetings\n\n| Aspect | Meeting | Podcast / Interview |\n|--------|---------|---------------------|\n| Speakers | 3–15+, often unknown | 2–3, usually known (host + guests) |\n| Language | Usually single | May mix languages (bilingual hosts) |\n| Hotwords | Participant names + terms | Show name, guest name, topic terms |\n| Speaker context | Role-based keywords | Host asks questions, guest answers |\n| Diarization | Critical | Easier (fewer, distinct voices) |\n\n### Recommended settings\n\n```bash\n# English podcast (2 speakers, host + guest)\npython3 transcribe_funasr.py episode.flac --lang en --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n\n# Bilingual podcast (auto-detect language switches)\npython3 transcribe_funasr.py episode.flac --lang auto --num-speakers 2 \\\n  --speakers \"Alice,Bob\"\n\n# Chinese podcast with topic hotwords\npython3 transcribe_funasr.py episode.flac --lang zh --num-speakers 3 \\\n  --speakers \"主持人,嘉宾A,嘉宾B\" \\\n  --hotwords \"播客名 嘉宾全名 讨论主题关键词\"\n\n# Multi-language podcast (e.g., Spanish + English)\npython3 transcribe_funasr.py episode.flac --lang whisper --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n```\n\n### Tips for podcast transcription\n\n1. **Always provide `--num-speakers`** — podcasts have a known, fixed speaker count;\n   this dramatically improves diarization accuracy with only 2–3 voices\n2. **Always provide `--speakers`** — host/guest names are known upfront\n3. **`--lang auto`** works for bilingual transcript-only output (no speaker labels) —\n   SenseVoiceSmall handles intra-utterance language switching (zh/en/ja/ko/yue) but\n   does not output timestamps, so **speaker diarization is not supported**.\n   Use `--lang zh` for Chinese podcasts that need speaker identification.\n4. **`--lang whisper`** for any other language or heavy code-switching (also lacks\n   timestamp support for diarization — transcript-only)\n5. **Hotwords** — for Chinese podcasts, include the show name and guest's full name;\n   for English podcasts, hotwords are usually unnecessary\n6. **`--speaker-context`** — describe the host/guest dynamic:\n   ```json\n   {\n     \"Alice\": \"Host, asks questions, introduces topics, wraps up segments\",\n     \"Bob\": \"Guest, expert on topic X, shares personal anecdotes\"\n   }\n   ```\n7. **Audio quality** — podcasts are typically studio-recorded with better SNR than\n   meetings; diarization accuracy is correspondingly higher\n8. **Audio source matters** — mobile app downloads may be truncated (trial/preview\n   versions). Download from the web interface for complete files. The script's\n   Phase 0 duration validation catches conversion truncation but cannot detect\n   a source file that is already incomplete.\n\n## `--lang mimo` — Xiaomi MiMo-V2.5-ASR (local, GPU)\n\nMiMo is an 8B-parameter LLM-based ASR model from Xiaomi. It outputs plain text\nwith no per-sentence timestamps and no speaker labels. The `funasr-transcribe`\nskill wraps it in a VAD + speaker-clustering sandwich so output format matches\nthe FunASR presets:\n\n```\nPhase 1a  FSMN VAD           → [(start_ms, end_ms), ...]\nPhase 1b  MiMo asr_sft()     → text per VAD segment\nPhase 1c  CAM++ + KMeans     → speaker ID per VAD segment\n```\n\nFiles: `scripts/mimo_asr.py` (orchestrator), `scripts/setup_mimo.sh` (installer).\n\n### Expected RTF\n\nOn a single A100 (40 GB), 4h audio → ~24 min wall clock for Phase 1 (RTF ≈\n0.1). Compare to `--lang zh` on the same GPU at RTF ≈ 0.02–0.05. Trade speed\nfor reported accuracy gains on dialects, code-switching, and lyrics.\n\n### Resume (`--resume-mimo`)\n\nSegment-level failures (OOM, CUDA error) retry 3× with\n[0.5s, 2s, 5s] backoff after `gc.collect()` + `torch.cuda.empty_cache()`. If all\nretries fail, a `*_mimo_partial.json` file captures VAD segments, completed\ntranscriptions, and the failed index. `--resume-mimo` picks up from the failed\nsegment, verifying audio SHA256 + `--mimo-audio-tag` match before continuing.\n\nArchive v1.6.0: 10 files, 60179 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (3814b), scripts/speaker_gender.py (11457b), scripts/test_speaker_verification.py (74555b), scripts/transcribe_funasr.py (60454b), scripts/verify_speakers.py (17584b), SKILL.md (13002b), _meta.json (150b)\n\nFile v1.6.0:SKILL.md\n\n---\nname: funasr-transcribe\nversion: 1.6.0\ndescription: >\n  This skill should be used when the user explicitly asks to \"transcribe a meeting\",\n  \"transcribe audio\", \"transcribe a meeting recording\",\n  \"convert audio to text\", \"generate meeting minutes from audio\",\n  \"do speech-to-text\", \"transcribe with speaker diarization\",\n  \"identify speakers in audio\", \"transcribe Chinese audio\",\n  \"transcribe English audio\", \"transcribe Japanese audio\",\n  \"multi-speaker transcription\", \"transcribe a podcast\",\n  \"transcribe podcast episode\", \"transcribe an interview\",\n  \"convert podcast to text\", \"podcast to transcript\",\n  or mentions FunASR, Paraformer, SenseVoice, Whisper, meeting\n  transcription, podcast transcription, or speaker diarization.\n  Supports multi-speaker meeting and podcast transcription in Chinese,\n  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper)\n  with automatic speaker diarization and hotword biasing.\n  Works on both GPU and CPU.\nmetadata:\n  openclaw:\n    requires:\n      bins: [\"python3\", \"ffmpeg\"]\n    env_vars:\n      - name: AWS_REGION\n        required: false\n        description: \"AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed.\"\n      - name: ANTHROPIC_API_KEY\n        required: false\n        description: \"API key for Anthropic Claude LLM cleanup\"\n      - name: OPENAI_API_KEY\n        required: false\n        description: \"API key for OpenAI-compatible LLM cleanup\"\n      - name: OPENAI_BASE_URL\n        required: false\n        description: \"Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)\"\n    emoji: \"🎙️\"\n    homepage: \"https://github.com/zxkane/audio-transcriber-funasr\"\n---\n\n# FunASR Meeting & Podcast Transcription\n\nTranscribe multi-speaker audio into structured Markdown with automatic\nspeaker diarization, hotword biasing, and optional LLM cleanup.\n\nAll scripts run directly from the plugin directory — no copying needed.\nDefine this shorthand at the start of every session:\n\n```bash\nSCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/funasr-transcribe/scripts\n```\n\n## Supported Languages\n\n| `--lang` | Model | Languages | Hotword |\n|----------|-------|-----------|---------|\n| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |\n| `zh-basic` | Paraformer-large | Chinese | No |\n| `en` | Paraformer-en | English | No |\n| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |\n| `whisper` | Whisper-large-v3-turbo | 99 languages | No |\n\nAll presets include **speaker diarization** (CAM++) and **VAD** (FSMN).\n\n> **Diarization caveat:** `auto` and `whisper` do not output per-sentence timestamps,\n> so speaker diarization does not work with these presets. Use `zh`, `zh-basic`, or\n> `en` when speaker identification is needed (e.g., podcasts, meetings).\n\n## Workflow\n\nBefore starting transcription, **always ask the user**:\n\n1. **Audio file** — path to the recording (required)\n2. **Type** — meeting, podcast, or interview? (affects defaults)\n3. **Language** — what language is spoken? (default: Chinese)\n4. **Number of speakers** — how many participants? (improves diarization)\n5. **Speaker names** — for podcasts: host + guest names; for meetings: attendee list\n6. **Supporting files** — ask:\n   > \"Do you have any of the following to improve accuracy?\"\n   > - **Attendee / guest list** — for hotwords and speaker mapping\n   > - **Meeting agenda or episode topic** — for hotwords (terms, names)\n   > - **Reference documents** (show notes, prior notes) — for speaker identification and ASR correction\n\n**Adapt defaults by recording type:**\n- **Meeting**: default `--lang zh`, ask about supporting files\n- **Podcast / interview**: default `--lang zh`, `--num-speakers 2`, always ask for\n  host + guest names, suggest `--speaker-context` for roles\n  (do NOT use `--lang auto` — it lacks timestamps for speaker diarization)\n\n> **⚠️ `--speakers` must use the speaker's real name, not a podcast alias.**\n> The value passed to `--speakers` is used verbatim as the speaker label in the\n> output transcript. Always derive it from the host/guest's actual name (e.g.\n> from a shownotes \"Host:\" field), not from the podcast feed name or title.\n>\n> Example: if shownotes lists \"Host: 张三（张三的播客）\", pass `--speakers '张三'`\n> — not the alias \"张三的播客\". Add both the real name and the alias to\n> `hotwords.txt` so ASR can recognise both forms.\n>\n> When both `--speakers` and `--reference` are supplied, the script detects\n> this mistake at startup and prints an `ACTION REQUIRED` block naming the\n> suggested real name. **If you see that block, stop the run and re-invoke\n> with the corrected `--speakers` value before Phase 3** — the warning does\n> not abort the pipeline.\n\nIf the user provides supporting materials:\n- Extract participant names and key terms → create `hotwords.txt` (include both real name and alias)\n- Extract per-person context → create `speaker-context.json`\n- Pass original reference document with `--reference`\n- Use all three together for best results\n\n## Quick Start\n\n### 1. Environment Setup\n\n```bash\nAUTO_YES=1 bash $SCRIPTS/setup_env.sh\n# Or force CPU:  AUTO_YES=1 bash $SCRIPTS/setup_env.sh cpu\n```\n\nThe setup script patches FunASR's spectral clustering for O(N²·k) performance.\nWithout this, recordings over ~1 hour hang for hours during speaker clustering.\n\n### 2. Run Transcription\n\nOutput files are written to the current working directory.\n\n**LLM cleanup (Phase 3) is opt-in.** By default, transcription runs locally\nwithout contacting any external service. To enable LLM-powered ASR correction\nand speaker name refinement, pass `--model <model-id>`. Use LLM cleanup when:\n- The raw transcript has many ASR errors (names, technical terms)\n- You need polished, publication-ready output\n- Speaker names need to be refined from context\n\n> **⚠️ Data Privacy:** When LLM cleanup is enabled via `--model`, transcript\n> excerpts are sent to external LLM providers (AWS Bedrock, Anthropic, or\n> OpenAI depending on the model ID). Use `--skip-llm` or omit `--model` to\n> keep all data local. For Bedrock, boto3 uses the standard AWS credential\n> chain (IAM role, SSO, `~/.aws/credentials`, env vars).\n\n```bash\n# Chinese meeting with hotwords (local-only, no LLM)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt\n\n# English meeting with speaker names\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang en --speakers \"Alice,Bob,Carol,Dave\"\n\n# Auto-detect language (zh/en/ja/ko/yue)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang auto --num-speakers 6\n\n# Whisper for any language\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang whisper --num-speakers 4\n\n# Enable LLM cleanup for polished output (requires --model)\n# Bedrock (uses AWS credential chain: IAM role, SSO, ~/.aws/credentials)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Anthropic API (requires ANTHROPIC_API_KEY env var)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --model claude-sonnet-4-6\n\n# OpenAI-compatible API (requires OPENAI_API_KEY env var)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --model gpt-4o\n\n# Full pipeline with all supporting files + LLM (best quality)\npython3 $SCRIPTS/transcribe_funasr.py episode.m4a \\\n    --lang zh --num-speakers 2 \\\n    --hotwords hotwords.txt \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json \\\n    --reference show-notes.md \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Resume interrupted LLM cleanup\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --skip-transcribe --model us.anthropic.claude-sonnet-4-6\n```\n\n### 3. Verify Speaker Labels\n\nIf the transcript has swapped speaker labels (common with podcasts),\nthe verification script can detect and fix mismatches using LLM analysis:\n\n```bash\n# Dry-run: check if host/guest are swapped\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json\n\n# Apply the fix\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json --fix\n\n# Multi-speaker meeting: full reassignment\npython3 $SCRIPTS/verify_speakers.py meeting_raw_transcript.json \\\n    --speakers \"Alice,Bob,Carol,Dave\" \\\n    --speaker-context speaker-context.json --fix\n\n# Then regenerate the markdown with corrected labels\npython3 $SCRIPTS/transcribe_funasr.py original.m4a \\\n    --skip-transcribe --clean-cache\n```\n\nThe script analyzes the first 5 minutes (configurable with `--minutes`)\nand auto-detects podcast (2 speakers, swap detection) vs meeting\n(N speakers, full reassignment).\n\n## Audio Preprocessing\n\nThe script automatically converts input audio to 16kHz mono FLAC and\nvalidates that no audio is lost (detects silent truncation).\n\n| Format | 4h14m meeting | Quality | Recommendation |\n|--------|--------------|---------|----------------|\n| **FLAC** | **219MB** | Lossless | **Default, safest** |\n| Opus | 55MB | Lossy | Risk of truncation on long files |\n| WAV | 465MB | Lossless | Works but larger |\n| Original M4A | 173MB | Source | Also works directly |\n\n**Do NOT split long recordings** — splitting breaks speaker ID consistency.\n\n## Key Flags\n\n| Flag | Purpose |\n|------|---------|\n| `--lang` | `zh` (default), `zh-basic`, `en`, `auto`, `whisper` |\n| `--hotwords` | Hotword file or string — biases ASR (zh only) |\n| `--reference F` | Reference file for LLM ASR correction |\n| `--num-speakers N` | Expected speaker count (improves diarization) |\n| `--speakers \"A,B,C\"` | Assign real names by first-appearance order |\n| `--speaker-context F` | JSON with per-speaker roles for LLM |\n| `--no-detect-gender` | Disable automatic speaker gender detection (CAM++ gender classifier) |\n| `--speaker-genders \"A:female,B:male\"` | Override per-speaker gender (also accepts positional `female,male`) |\n| `--audio-format` | `flac` (default), `opus`, `wav` |\n| `--device cpu` | Force CPU mode |\n| `--batch-size N` | Adjust for memory (60 for CPU, 100 if GPU OOM) |\n| `--phase1-only` | Exit after Phase 1 (VAD + ASR + diarization), skip Phase 2 + 3 |\n| `--json-out PATH` | Write raw transcript JSON to explicit path (overrides default naming) |\n| `--skip-transcribe` | Resume from saved `*_raw_transcript.json` |\n| `--skip-llm` | Skip LLM cleanup (default when `--model` is omitted) |\n| `--model ID` | Enable LLM cleanup with this model (auto-detects Bedrock/Anthropic/OpenAI) |\n| `--title \"...\"` | Output document title |\n| `--clean-cache` | Delete LLM chunk cache after completion |\n| `--output PATH` | Custom output file path |\n| `--model-cache-dir` | ModelScope model cache directory (~3GB, default: `~/.cache/modelscope/`) |\n\n## Outputs\n\n- `<stem>-transcript.md` — Final Markdown with speaker labels and timestamps\n- `<stem>_raw_transcript.json` — Raw Phase 1 output (for resume/analysis)\n\n## Speaker Diarization Tips\n\nFunASR's CAM++ may merge acoustically similar speakers. To improve:\n\n1. **`--num-speakers N`** — Hint expected count\n2. **`--hotwords`** — Include participant names (Chinese names work best)\n3. **`--speaker-context`** — Provide per-person keywords for LLM splitting\n4. **Keyword matching** — Search `*_raw_transcript.json` for unique phrases\n\n### Speaker gender\n\nEnabled by default: each detected speaker is classified as `male` / `female`\nvia 3D-Speaker's CAM++ gender classifier (`iic/speech_campplus_two_class_gender_16k`).\nThe result appears next to each name in the **Speaker List** table and is\ninjected into the LLM cleanup prompt so pronouns (他/她, he/she) get corrected.\n\nPrecedence when combined:\n\n1. `--speaker-genders \"Alice:female,Bob:male\"` (explicit CLI) — always wins\n2. Reference text hints like `主播（女）：韩梅梅` or `Host (male): Alice` — override auto\n3. CAM++ auto-detection — fallback\n\nDisable with `--no-detect-gender` if you don't need gender and want to save\nthe ~500 MB model download and extra inference time.\n\n## CPU-only / Low-Memory Machines\n\nLong recordings on resource-constrained machines may hit exec timeouts\nor OOM kills. See `references/pipeline-details.md` for workarounds:\n- Detach from agent timeouts with `systemd-run` or `nohup`\n- Prevent OOM via swap and/or `--lang zh-basic` (lighter model)\n\n## Additional Resources\n\n- **`references/pipeline-details.md`** — Architecture, model specs, benchmarks,\n  speaker role verification, hotword effectiveness, clustering patch\n- **`scripts/transcribe_funasr.py`** — Main transcription pipeline\n- **`scripts/verify_speakers.py`** — Speaker label verification & fix\n- **`scripts/llm_utils.py`** — Shared LLM infrastructure (Bedrock/Anthropic/OpenAI)\n- **`scripts/setup_env.sh`** — Environment setup (venv + deps + patch)\n\nFile v1.6.0:_meta.json\n\n{\n  \"ownerId\": \"kn72agp4n0v3y1gk4qds89r0wn82msn3\",\n  \"slug\": \"zxkane-audio-transcriber-funasr\",\n  \"version\": \"1.6.0\",\n  \"publishedAt\": 1777374401390\n}\n\nFile v1.6.0:references/pipeline-details.md\n\n# FunASR Meeting Transcription Pipeline — Technical Details\n\n## Architecture\n\n```\nAudio File (.m4a/.mp3/.wav)\n  │\n  ├─ [ffmpeg] ──► 16kHz mono WAV\n  │\n  ├─ [Phase 1: FunASR] ──► raw_transcript.json\n  │   ├─ FSMN-VAD: segment speech vs silence\n  │   ├─ ASR model (language-dependent, see below)\n  │   ├─ (Optional) Hotword biasing (SeACo-Paraformer only)\n  │   ├─ Punctuation restoration (model-dependent)\n  │   └─ CAM++: speaker embeddings → spectral clustering\n  │\n  ├─ [Phase 2: Post-process]\n  │   ├─ Merge consecutive same-speaker utterances (<2s gap)\n  │   ├─ Map speaker IDs to names (if provided)\n  │   └─ Auto-verify via self-introduction detection\n  │\n  └─ [Phase 3: LLM cleanup] ──► transcript.md\n      ├─ LLM speaker role verification (if --speaker-context provided)\n      ├─ Remove fillers (um, uh, 嗯, 啊, etc.)\n      ├─ Fix ASR errors (homophones, context-based)\n      ├─ Polish grammar while preserving meaning\n      └─ (Optional) Identify merged speakers via context\n```\n\n## Language Presets & Models\n\n### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |\n| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |\n| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |\n| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |\n\nSeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)\nto bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.\n\n### `--lang zh-basic` (Chinese, no hotword)\n\nSame as `zh` but uses the base Paraformer-large without hotword support.\nUse when hotword biasing is unnecessary or causing issues with English terms.\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |\n\n### `--lang en` (English)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |\n\n### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |\n\nSenseVoiceSmall includes built-in punctuation and supports emotion detection.\n\n### `--lang whisper` (Multilingual, 99 languages)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |\n\nAll presets share the same VAD (`fsmn-vad`) and speaker diarization (`cam++`) models.\nModels are auto-downloaded from ModelScope on first run.\n\n## Hotword Biasing\n\nSeACo-Paraformer (`--lang zh`) supports hotword customization to improve recognition\nof specific terms — particularly useful for participant names, project names, and\ndomain-specific jargon in meetings.\n\n### How to provide hotwords\n\n```bash\n# Space-separated string\npython3 transcribe_funasr.py meeting.wav --lang zh --hotwords \"张三 李四 ClawCon Rebase\"\n\n# Text file (one word per line)\npython3 transcribe_funasr.py meeting.wav --lang zh --hotwords hotwords.txt\n```\n\n### What to include in hotwords\n\nFor meeting transcription, a good hotwords file includes:\n- **Participant names** (full names in the meeting's language)\n- **Project / product names** (internal codenames, product brands)\n- **Domain-specific Chinese terms** (technical jargon, acronyms in Chinese)\n- **Organization names** (company, team, department names)\n\n### Effectiveness — empirical results\n\nTested on a 4h14m, 9-speaker Chinese meeting with 27 hotwords:\n\n| Term | Without hotwords | With hotwords | Change |\n|------|-----------------|---------------|--------|\n| 龙虾 (lobechat) | 28 | 42 | **+50%** |\n| 高琦 (person name) | 0 | 7 | **0 → 7** |\n| 搬瓦工 (BandwagonHost) | 0 | 1 | **0 → 1** |\n| 谢锐 (person name) | 0 | 1 | improved |\n| 鲲鹏 (org name) | 6 | 7 | slight improvement |\n| Rebase (English) | 5 | 0 | **regression** |\n| Tailwind (English) | 3 | 1 | **regression** |\n\n**Key findings:**\n1. **Chinese terms benefit significantly** — names, brands, and Chinese jargon\n   see clear improvement (龙虾 +50%, 高琦 from zero)\n2. **English terms may regress** — SeACo's hotword biasing operates on Chinese\n   token vocabulary; English loanwords can be disrupted\n3. **Person names have limited uplift** — meeting participants rarely say full\n   names aloud; hotwords help only when names do appear in speech\n\n**Recommendation:** Include Chinese terms and names in hotwords. For English\ntechnical terms, rely on Phase 3 LLM cleanup rather than hotword biasing.\nIf English term accuracy is critical and hotword biasing causes regressions,\nuse `--lang zh-basic` instead.\n\n## Performance Benchmarks\n\nTested on a 4h14m, 9-speaker Chinese meeting recording (GPU: L40S 46GB):\n\n| Metric | Paraformer-large | SeACo + Hotword | Notes |\n|--------|-----------------|-----------------|-------|\n| Model load | 14s | 14s (422s first run, downloading 944MB model) | Cached after first run |\n| Transcription | 169s | 168s | Virtually identical |\n| Raw sentences | 6672 | 6725 | Comparable |\n| Merged segments | 1695 | 1724 | Comparable |\n| Speakers detected | 7 (of 9) | 7 (of 9) | Same diarization result |\n\n| Metric | GPU (L40S 46GB) | CPU (estimated) |\n|--------|-----------------|-----------------|\n| Model load | 14s | ~30s |\n| Transcription | 169s | ~30-60 min |\n| Speaker clustering | ~10s (patched) | ~2-5 min (patched) |\n| LLM cleanup (17 chunks) | ~35 min | ~35 min (network-bound) |\n| Total | ~38 min | ~70-100 min |\n\n**Without the clustering patch**, the original `scipy.linalg.eigh()` on the full Laplacian\nmatrix was O(N^3) and took **10+ hours** on this recording. The patch reduces it to O(N^2*k)\nvia `scipy.sparse.linalg.eigsh()`.\n\n## Clustering Patch (Critical for Long Meetings)\n\nFunASR's `SpectralCluster.get_spec_embs()` uses `scipy.linalg.eigh(L)` which computes\nALL eigenvalues of the NxN Laplacian. For a 4-hour recording, N can be 6000+, making\nthis O(N^3) operation take hours.\n\nThe patch (`scripts/patch_clustering.py`) replaces this with:\n- `scipy.sparse.linalg.eigsh(L_sparse, k=num_speakers, which='SM')` — only computes\n  the k smallest eigenvalues needed, reducing complexity to O(N^2 * k)\n- Vectorized `p_pruning()` — replaces Python loop with numpy broadcasting\n\n**Always run the patch before processing meetings longer than ~1 hour.**\n\n## Speaker Role Verification\n\nSpeaker names are assigned by first-appearance order in the audio, which may\nswap host/guest labels (especially in podcasts). Two layers of automatic\nverification, plus a standalone post-hoc tool:\n\n1. **Phase 2 — self-introduction detection**: scans the first 5 minutes for\n   explicit self-introductions (\"我是X\", \"I'm X\") and swaps labels if mismatched\n2. **Phase 3 — LLM role verification**: when `--speaker-context` is provided,\n   the LLM analyzes the first chunk (up to 15 minutes) before cleanup begins.\n   For 2 speakers: binary CORRECT/SWAP detection. For 3+ speakers: full\n   JSON-based reassignment matching each label to the correct person.\n3. **Post-hoc — `verify_speakers.py`**: standalone script that verifies any\n   existing `*_raw_transcript.json`. Same two modes (2-speaker swap, N-speaker\n   reassignment) with dry-run support. See SKILL.md § Verify Speaker Labels.\n\nFor podcasts, always provide `--speaker-context` describing host/guest roles.\n\n## Speaker Diarization Limitations\n\nFunASR's CAM++ speaker diarization may merge acoustically similar speakers into one ID.\nIn the tested 9-person meeting, only 7 unique IDs were detected (two pairs merged).\n\nWorkarounds:\n1. **Provide `--num-speakers N`** to hint expected count (uses `preset_spk_num`)\n2. **Post-hoc keyword matching**: use reference documents (meeting agendas, attendee notes)\n   to identify which speaker ID maps to which person\n3. **LLM-assisted splitting**: provide `--speaker-context` with per-person keywords;\n   the LLM can then split merged speakers when context is clear (~73% success rate)\n\n## Supporting Files for Better Results\n\nPrepare these files before transcription for best results:\n\n| File | Used in | Purpose |\n|------|---------|---------|\n| `hotwords.txt` | Phase 1 (`--hotwords`) | Bias ASR toward names and terms |\n| `speaker-context.json` | Phase 3 (`--speaker-context`) | Help LLM identify and split speakers |\n| Meeting agenda | Manual reference | Identify meeting phases for post-analysis |\n| Attendee list | Build hotwords + speaker names | Map speaker IDs to real names |\n\n### Example: preparing supporting files from a meeting invite\n\n```bash\n# 1. Create hotwords.txt from attendee list and agenda\ncat > hotwords.txt << 'EOF'\nAlice\nBob\nCarol\nProjectAlpha\nSprint Review\nQ2 OKR\nEOF\n\n# 2. Create speaker-context.json from attendee roles\ncat > speaker-context.json << 'EOF'\n{\n  \"Alice\": \"Engineering manager, discusses sprint velocity and tech debt\",\n  \"Bob\": \"Product manager, presents roadmap and customer feedback\",\n  \"Carol\": \"Designer, shows mockups, mentions Figma and user testing\"\n}\nEOF\n\n# 3. Run with both\npython3 transcribe_funasr.py meeting.wav \\\n  --lang zh --num-speakers 3 \\\n  --speakers \"Alice,Bob,Carol\" \\\n  --hotwords hotwords.txt \\\n  --speaker-context speaker-context.json\n```\n\n## Audio Preprocessing\n\nFunASR works best with 16kHz mono audio. **FLAC is recommended** over WAV — lossless\nquality at ~50% the file size, and FunASR reads it natively via soundfile.\n\n```bash\n# Recommended: FLAC (lossless, compact)\nffmpeg -i recording.m4a -ar 16000 -ac 1 -sample_fmt s16 meeting.flac\n\n# Alternative: WAV (lossless, larger)\nffmpeg -i recording.m4a -ar 16000 -ac 1 meeting.wav\n```\n\n**Important:** Use `-sample_fmt s16` when converting to FLAC — without it, ffmpeg\nmay output 24-bit samples (s32/24bit) which doubles the file size with no ASR benefit.\n\n### Format comparison (4h14m meeting)\n\n| Format | Size | Quality | FunASR support |\n|--------|------|---------|---------------|\n| **FLAC (16kHz mono s16)** | **219MB** | Lossless | Native (soundfile) |\n| WAV (16kHz mono) | 465MB | Lossless | Native (soundfile) |\n| Opus (32kbps) | 54MB | Lossy | Native (soundfile) |\n| M4A/AAC (original 48kHz) | 173MB | Source | Via librosa |\n| M4A/AAC (16kHz 32kbps) | 60MB | Lossy | Via librosa |\n\nFunASR accepts all common audio formats. FLAC offers the best trade-off: lossless\nquality, reasonable size, and native reader support without librosa fallback.\n\nFor long recordings, do NOT split the audio — FunASR handles arbitrarily long files\nand splitting breaks speaker consistency across segments.\n\n## Resume / Checkpoint Support\n\nThe pipeline supports resuming interrupted runs:\n- **Phase 1 output**: `<stem>_raw_transcript.json` — use `--skip-transcribe` to skip ASR\n- **Phase 3 cache**: `<stem>_llm_cache/chunk_NNN.txt` — already-cleaned chunks are reused\n  (kept by default; add `--clean-cache` to delete after completion)\n\n## Model Caching\n\nFunASR models (~3 GB for the `zh` preset) are downloaded from ModelScope on first run\nand cached in `~/.cache/modelscope/hub/`. On ephemeral instances (EC2, cloud VMs), the\ncache is lost when the instance is replaced, requiring a ~2 minute re-download.\n\nTo persist the cache on durable storage (e.g., an EBS data volume):\n\n```bash\n# Via CLI flag (recommended)\npython3 transcribe_funasr.py meeting.flac --model-cache-dir /data/modelscope-cache ...\n\n# Via environment variable\nMODELSCOPE_CACHE=/data/modelscope-cache python3 transcribe_funasr.py meeting.flac ...\n```\n\nThe `systemd-run` examples below include `-E MODELSCOPE_CACHE=...` for this reason.\n\n## Speaker Context JSON Format\n\nThe `--speaker-context` file helps the LLM identify speakers and fix ASR errors:\n\n```json\n{\n  \"Alice\": \"Discussed Q1 revenue targets, mentioned Chicago office relocation\",\n  \"Bob\": \"Presented the new CI/CD pipeline, uses Terraform and ArgoCD\",\n  \"Carol\": \"HR updates, mentioned hiring freeze and new PTO policy\"\n}\n```\n\nThe context is injected into the LLM system prompt for each cleanup chunk.\n\n## Running on CPU-only / Low-Memory Machines\n\nLong recordings (2+ hours) on resource-constrained machines (CPU-only, ≤8 GB RAM)\nface two common failure modes. Both are silent — the process is killed mid-run\nwith no output files saved.\n\n### Problem 1: Process killed by execution timeout\n\nAI coding agents (Claude Code, OpenClaw, Cursor, etc.) impose execution timeouts\non shell commands — typically 2–10 minutes. On a 4-hour recording, CPU transcription\ntakes 1.5–2 hours, well past any agent timeout. The process is silently killed.\n\n**Fix — detach the ASR phase from the agent's process supervision.** Use `--skip-llm`\nfor the detached run because Phase 1 (ASR) is the CPU-intensive bottleneck; Phase 3\n(LLM cleanup) is network-bound and fast — run it afterward via `--skip-transcribe`\nunder the normal agent session.\n\nOption A: `systemd-run` (preferred on systemd hosts):\n\n```bash\nsystemd-run --user --unit=transcribe-job \\\n  --working-directory=/tmp \\\n  -E MODELSCOPE_CACHE=/data/modelscope-cache \\\n  bash -c 'source /path/to/.venv/bin/activate && \\\n    python3 /path/to/transcribe_funasr.py /tmp/meeting.flac \\\n    --lang zh --num-speakers 9 --skip-llm > /tmp/transcribe.log 2>&1'\n\n# Monitor progress\nsystemctl --user status transcribe-job.service\ntail -f /tmp/transcribe.log\n\n# Check result\nls -lh /tmp/*-transcript.md /tmp/*_raw_transcript.json\n```\n\n`systemd-run` creates a transient systemd service fully independent of the agent\nsession — it survives session resets, context pruning, and exec timeouts.\n\n> **Warning:** `systemd-run` creates an isolated mount namespace. FUSE mounts\n> (rclone, sshfs, Google Drive, etc.) from the parent session are NOT visible\n> to the transient service. Copy all dependency files (audio, hotwords,\n> speaker-context, reference documents) to a local path (e.g., `/tmp`) before\n> launching. Always use `--working-directory=/tmp` or another real filesystem path.\n\n> **Note:** `systemd-run --user` requires a user-level systemd instance. It may not\n> work in Docker containers or cloud VMs without `loginctl enable-linger`.\n\nOption B: `nohup` (works everywhere):\n\n```bash\nnohup bash -c 'source .venv/bin/activate && python3 transcribe_funasr.py meeting.flac \\\n  --lang zh --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\necho $!  # Save PID for monitoring\ntail -f transcribe.log\n```\n\n### Problem 2: OOM kill on machines with ≤8 GB RAM\n\nThe `zh` preset loads 4 model components simultaneously (SeACo-Paraformer + VAD +\nPunctuation + CAM++ speaker). Peak RSS can exceed 7 GB on a 4-hour recording. On\nmachines without swap, the OOM killer terminates the process silently.\n\n**Fix A — add swap before running** (requires root):\n\n```bash\nsudo fallocate -l 4G /swapfile\nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n\n# After transcription, optionally remove swap\nsudo swapoff /swapfile && sudo rm /swapfile\n```\n\n**Fix B — use `zh-basic` instead of `zh`:**\n\n`zh-basic` (Paraformer-large) loads one fewer model component than `zh`\n(SeACo-Paraformer), reducing peak RSS by ~1–1.5 GB. Accuracy is slightly lower\n(no hotword biasing) but sufficient for most meetings:\n\n```bash\npython3 transcribe_funasr.py meeting.flac --lang zh-basic --num-speakers 9 --skip-llm\n```\n\n**Combining both fixes** (swap + `zh-basic`) reliably handles 4+ hour recordings on\nmachines with as little as 8 GB RAM + 4 GB swap.\n\n### Recommended CPU workflow for long recordings\n\n```bash\n# 1. Add swap if RAM ≤ 8 GB\nsudo fallocate -l 4G /swapfile && sudo chmod 600 /swapfile \\\n  && sudo mkswap /swapfile && sudo swapon /swapfile\n\n# 2. Launch transcription detached from agent timeout\nnohup bash -c 'source .venv/bin/activate && python3 transcribe_funasr.py meeting.flac \\\n  --lang zh-basic --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\n# 3. Monitor\ntail -f transcribe.log\n\n# 4. When done, resume with LLM cleanup (network-bound, runs fine under agent)\npython3 transcribe_funasr.py meeting.flac --skip-transcribe\n```\n\n## Podcast Transcription\n\nThe pipeline handles podcasts and interviews with the same engine, but the workflow\ndiffers from meetings:\n\n### Key differences from meetings\n\n| Aspect | Meeting | Podcast / Interview |\n|--------|---------|---------------------|\n| Speakers | 3–15+, often unknown | 2–3, usually known (host + guests) |\n| Language | Usually single | May mix languages (bilingual hosts) |\n| Hotwords | Participant names + terms | Show name, guest name, topic terms |\n| Speaker context | Role-based keywords | Host asks questions, guest answers |\n| Diarization | Critical | Easier (fewer, distinct voices) |\n\n### Recommended settings\n\n```bash\n# English podcast (2 speakers, host + guest)\npython3 transcribe_funasr.py episode.flac --lang en --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n\n# Bilingual podcast (auto-detect language switches)\npython3 transcribe_funasr.py episode.flac --lang auto --num-speakers 2 \\\n  --speakers \"Alice,Bob\"\n\n# Chinese podcast with topic hotwords\npython3 transcribe_funasr.py episode.flac --lang zh --num-speakers 3 \\\n  --speakers \"主持人,嘉宾A,嘉宾B\" \\\n  --hotwords \"播客名 嘉宾全名 讨论主题关键词\"\n\n# Multi-language podcast (e.g., Spanish + English)\npython3 transcribe_funasr.py episode.flac --lang whisper --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n```\n\n### Tips for podcast transcription\n\n1. **Always provide `--num-speakers`** — podcasts have a known, fixed speaker count;\n   this dramatically improves diarization accuracy with only 2–3 voices\n2. **Always provide `--speakers`** — host/guest names are known upfront\n3. **`--lang auto`** works for bilingual transcript-only output (no speaker labels) —\n   SenseVoiceSmall handles intra-utterance language switching (zh/en/ja/ko/yue) but\n   does not output timestamps, so **speaker diarization is not supported**.\n   Use `--lang zh` for Chinese podcasts that need speaker identification.\n4. **`--lang whisper`** for any other language or heavy code-switching (also lacks\n   timestamp support for diarization — transcript-only)\n5. **Hotwords** — for Chinese podcasts, include the show name and guest's full name;\n   for English podcasts, hotwords are usually unnecessary\n6. **`--speaker-context`** — describe the host/guest dynamic:\n   ```json\n   {\n     \"Alice\": \"Host, asks questions, introduces topics, wraps up segments\",\n     \"Bob\": \"Guest, expert on topic X, shares personal anecdotes\"\n   }\n   ```\n7. **Audio quality** — podcasts are typically studio-recorded with better SNR than\n   meetings; diarization accuracy is correspondingly higher\n8. **Audio source matters** — mobile app downloads may be truncated (trial/preview\n   versions). Download from the web interface for complete files. The script's\n   Phase 0 duration validation catches conversion truncation but cannot detect\n   a source file that is already incomplete.\n\nArchive v1.5.1: 9 files, 52135 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (3814b), scripts/test_speaker_verification.py (60603b), scripts/transcribe_funasr.py (57317b), scripts/verify_speakers.py (17584b), SKILL.md (12079b), _meta.json (150b)\n\nFile v1.5.1:SKILL.md\n\n---\nname: funasr-transcribe\nversion: 1.5.1\ndescription: >\n  This skill should be used when the user explicitly asks to \"transcribe a meeting\",\n  \"transcribe audio\", \"transcribe a meeting recording\",\n  \"convert audio to text\", \"generate meeting minutes from audio\",\n  \"do speech-to-text\", \"transcribe with speaker diarization\",\n  \"identify speakers in audio\", \"transcribe Chinese audio\",\n  \"transcribe English audio\", \"transcribe Japanese audio\",\n  \"multi-speaker transcription\", \"transcribe a podcast\",\n  \"transcribe podcast episode\", \"transcribe an interview\",\n  \"convert podcast to text\", \"podcast to transcript\",\n  or mentions FunASR, Paraformer, SenseVoice, Whisper, meeting\n  transcription, podcast transcription, or speaker diarization.\n  Supports multi-speaker meeting and podcast transcription in Chinese,\n  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper)\n  with automatic speaker diarization and hotword biasing.\n  Works on both GPU and CPU.\nmetadata:\n  openclaw:\n    requires:\n      bins: [\"python3\", \"ffmpeg\"]\n    env_vars:\n      - name: AWS_REGION\n        required: false\n        description: \"AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed.\"\n      - name: ANTHROPIC_API_KEY\n        required: false\n        description: \"API key for Anthropic Claude LLM cleanup\"\n      - name: OPENAI_API_KEY\n        required: false\n        description: \"API key for OpenAI-compatible LLM cleanup\"\n      - name: OPENAI_BASE_URL\n        required: false\n        description: \"Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)\"\n    emoji: \"🎙️\"\n    homepage: \"https://github.com/zxkane/audio-transcriber-funasr\"\n---\n\n# FunASR Meeting & Podcast Transcription\n\nTranscribe multi-speaker audio into structured Markdown with automatic\nspeaker diarization, hotword biasing, and optional LLM cleanup.\n\nAll scripts run directly from the plugin directory — no copying needed.\nDefine this shorthand at the start of every session:\n\n```bash\nSCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/funasr-transcribe/scripts\n```\n\n## Supported Languages\n\n| `--lang` | Model | Languages | Hotword |\n|----------|-------|-----------|---------|\n| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |\n| `zh-basic` | Paraformer-large | Chinese | No |\n| `en` | Paraformer-en | English | No |\n| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |\n| `whisper` | Whisper-large-v3-turbo | 99 languages | No |\n\nAll presets include **speaker diarization** (CAM++) and **VAD** (FSMN).\n\n> **Diarization caveat:** `auto` and `whisper` do not output per-sentence timestamps,\n> so speaker diarization does not work with these presets. Use `zh`, `zh-basic`, or\n> `en` when speaker identification is needed (e.g., podcasts, meetings).\n\n## Workflow\n\nBefore starting transcription, **always ask the user**:\n\n1. **Audio file** — path to the recording (required)\n2. **Type** — meeting, podcast, or interview? (affects defaults)\n3. **Language** — what language is spoken? (default: Chinese)\n4. **Number of speakers** — how many participants? (improves diarization)\n5. **Speaker names** — for podcasts: host + guest names; for meetings: attendee list\n6. **Supporting files** — ask:\n   > \"Do you have any of the following to improve accuracy?\"\n   > - **Attendee / guest list** — for hotwords and speaker mapping\n   > - **Meeting agenda or episode topic** — for hotwords (terms, names)\n   > - **Reference documents** (show notes, prior notes) — for speaker identification and ASR correction\n\n**Adapt defaults by recording type:**\n- **Meeting**: default `--lang zh`, ask about supporting files\n- **Podcast / interview**: default `--lang zh`, `--num-speakers 2`, always ask for\n  host + guest names, suggest `--speaker-context` for roles\n  (do NOT use `--lang auto` — it lacks timestamps for speaker diarization)\n\n> **⚠️ `--speakers` must use the speaker's real name, not a podcast alias.**\n> The value passed to `--speakers` is used verbatim as the speaker label in the\n> output transcript. Always derive it from the host/guest's actual name (e.g.\n> from a shownotes \"Host:\" field), not from the podcast feed name or title.\n>\n> Example: if shownotes lists \"Host: 张三（张三的播客）\", pass `--speakers '张三'`\n> — not the alias \"张三的播客\". Add both the real name and the alias to\n> `hotwords.txt` so ASR can recognise both forms.\n>\n> When both `--speakers` and `--reference` are supplied, the script detects\n> this mistake at startup and prints an `ACTION REQUIRED` block naming the\n> suggested real name. **If you see that block, stop the run and re-invoke\n> with the corrected `--speakers` value before Phase 3** — the warning does\n> not abort the pipeline.\n\nIf the user provides supporting materials:\n- Extract participant names and key terms → create `hotwords.txt` (include both real name and alias)\n- Extract per-person context → create `speaker-context.json`\n- Pass original reference document with `--reference`\n- Use all three together for best results\n\n## Quick Start\n\n### 1. Environment Setup\n\n```bash\nAUTO_YES=1 bash $SCRIPTS/setup_env.sh\n# Or force CPU:  AUTO_YES=1 bash $SCRIPTS/setup_env.sh cpu\n```\n\nThe setup script patches FunASR's spectral clustering for O(N²·k) performance.\nWithout this, recordings over ~1 hour hang for hours during speaker clustering.\n\n### 2. Run Transcription\n\nOutput files are written to the current working directory.\n\n**LLM cleanup (Phase 3) is opt-in.** By default, transcription runs locally\nwithout contacting any external service. To enable LLM-powered ASR correction\nand speaker name refinement, pass `--model <model-id>`. Use LLM cleanup when:\n- The raw transcript has many ASR errors (names, technical terms)\n- You need polished, publication-ready output\n- Speaker names need to be refined from context\n\n> **⚠️ Data Privacy:** When LLM cleanup is enabled via `--model`, transcript\n> excerpts are sent to external LLM providers (AWS Bedrock, Anthropic, or\n> OpenAI depending on the model ID). Use `--skip-llm` or omit `--model` to\n> keep all data local. For Bedrock, boto3 uses the standard AWS credential\n> chain (IAM role, SSO, `~/.aws/credentials`, env vars).\n\n```bash\n# Chinese meeting with hotwords (local-only, no LLM)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt\n\n# English meeting with speaker names\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang en --speakers \"Alice,Bob,Carol,Dave\"\n\n# Auto-detect language (zh/en/ja/ko/yue)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang auto --num-speakers 6\n\n# Whisper for any language\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang whisper --num-speakers 4\n\n# Enable LLM cleanup for polished output (requires --model)\n# Bedrock (uses AWS credential chain: IAM role, SSO, ~/.aws/credentials)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Anthropic API (requires ANTHROPIC_API_KEY env var)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --model claude-sonnet-4-6\n\n# OpenAI-compatible API (requires OPENAI_API_KEY env var)\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --model gpt-4o\n\n# Full pipeline with all supporting files + LLM (best quality)\npython3 $SCRIPTS/transcribe_funasr.py episode.m4a \\\n    --lang zh --num-speakers 2 \\\n    --hotwords hotwords.txt \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json \\\n    --reference show-notes.md \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Resume interrupted LLM cleanup\npython3 $SCRIPTS/transcribe_funasr.py meeting.wav \\\n    --skip-transcribe --model us.anthropic.claude-sonnet-4-6\n```\n\n### 3. Verify Speaker Labels\n\nIf the transcript has swapped speaker labels (common with podcasts),\nthe verification script can detect and fix mismatches using LLM analysis:\n\n```bash\n# Dry-run: check if host/guest are swapped\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json\n\n# Apply the fix\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json --fix\n\n# Multi-speaker meeting: full reassignment\npython3 $SCRIPTS/verify_speakers.py meeting_raw_transcript.json \\\n    --speakers \"Alice,Bob,Carol,Dave\" \\\n    --speaker-context speaker-context.json --fix\n\n# Then regenerate the markdown with corrected labels\npython3 $SCRIPTS/transcribe_funasr.py original.m4a \\\n    --skip-transcribe --clean-cache\n```\n\nThe script analyzes the first 5 minutes (configurable with `--minutes`)\nand auto-detects podcast (2 speakers, swap detection) vs meeting\n(N speakers, full reassignment).\n\n## Audio Preprocessing\n\nThe script automatically converts input audio to 16kHz mono FLAC and\nvalidates that no audio is lost (detects silent truncation).\n\n| Format | 4h14m meeting | Quality | Recommendation |\n|--------|--------------|---------|----------------|\n| **FLAC** | **219MB** | Lossless | **Default, safest** |\n| Opus | 55MB | Lossy | Risk of truncation on long files |\n| WAV | 465MB | Lossless | Works but larger |\n| Original M4A | 173MB | Source | Also works directly |\n\n**Do NOT split long recordings** — splitting breaks speaker ID consistency.\n\n## Key Flags\n\n| Flag | Purpose |\n|------|---------|\n| `--lang` | `zh` (default), `zh-basic`, `en`, `auto`, `whisper` |\n| `--hotwords` | Hotword file or string — biases ASR (zh only) |\n| `--reference F` | Reference file for LLM ASR correction |\n| `--num-speakers N` | Expected speaker count (improves diarization) |\n| `--speakers \"A,B,C\"` | Assign real names by first-appearance order |\n| `--speaker-context F` | JSON with per-speaker roles for LLM |\n| `--audio-format` | `flac` (default), `opus`, `wav` |\n| `--device cpu` | Force CPU mode |\n| `--batch-size N` | Adjust for memory (60 for CPU, 100 if GPU OOM) |\n| `--phase1-only` | Exit after Phase 1 (VAD + ASR + diarization), skip Phase 2 + 3 |\n| `--json-out PATH` | Write raw transcript JSON to explicit path (overrides default naming) |\n| `--skip-transcribe` | Resume from saved `*_raw_transcript.json` |\n| `--skip-llm` | Skip LLM cleanup (default when `--model` is omitted) |\n| `--model ID` | Enable LLM cleanup with this model (auto-detects Bedrock/Anthropic/OpenAI) |\n| `--title \"...\"` | Output document title |\n| `--clean-cache` | Delete LLM chunk cache after completion |\n| `--output PATH` | Custom output file path |\n| `--model-cache-dir` | ModelScope model cache directory (~3GB, default: `~/.cache/modelscope/`) |\n\n## Outputs\n\n- `<stem>-transcript.md` — Final Markdown with speaker labels and timestamps\n- `<stem>_raw_transcript.json` — Raw Phase 1 output (for resume/analysis)\n\n## Speaker Diarization Tips\n\nFunASR's CAM++ may merge acoustically similar speakers. To improve:\n\n1. **`--num-speakers N`** — Hint expected count\n2. **`--hotwords`** — Include participant names (Chinese names work best)\n3. **`--speaker-context`** — Provide per-person keywords for LLM splitting\n4. **Keyword matching** — Search `*_raw_transcript.json` for unique phrases\n\n## CPU-only / Low-Memory Machines\n\nLong recordings on resource-constrained machines may hit exec timeouts\nor OOM kills. See `references/pipeline-details.md` for workarounds:\n- Detach from agent timeouts with `systemd-run` or `nohup`\n- Prevent OOM via swap and/or `--lang zh-basic` (lighter model)\n\n## Additional Resources\n\n- **`references/pipeline-details.md`** — Architecture, model specs, benchmarks,\n  speaker role verification, hotword effectiveness, clustering patch\n- **`scripts/transcribe_funasr.py`** — Main transcription pipeline\n- **`scripts/verify_speakers.py`** — Speaker label verification & fix\n- **`scripts/llm_utils.py`** — Shared LLM infrastructure (Bedrock/Anthropic/OpenAI)\n- **`scripts/setup_env.sh`** — Environment setup (venv + deps + patch)\n\nFile v1.5.1:_meta.json\n\n{\n  \"ownerId\": \"kn72agp4n0v3y1gk4qds89r0wn82msn3\",\n  \"slug\": \"zxkane-audio-transcriber-funasr\",\n  \"version\": \"1.5.1\",\n  \"publishedAt\": 1777165891256\n}\n\nFile v1.5.1:references/pipeline-details.md\n\n# FunASR Meeting Transcription Pipeline — Technical Details\n\n## Architecture\n\n```\nAudio File (.m4a/.mp3/.wav)\n  │\n  ├─ [ffmpeg] ──► 16kHz mono WAV\n  │\n  ├─ [Phase 1: FunASR] ──► raw_transcript.json\n  │   ├─ FSMN-VAD: segment speech vs silence\n  │   ├─ ASR model (language-dependent, see below)\n  │   ├─ (Optional) Hotword biasing (SeACo-Paraformer only)\n  │   ├─ Punctuation restoration (model-dependent)\n  │   └─ CAM++: speaker embeddings → spectral clustering\n  │\n  ├─ [Phase 2: Post-process]\n  │   ├─ Merge consecutive same-speaker utterances (<2s gap)\n  │   ├─ Map speaker IDs to names (if provided)\n  │   └─ Auto-verify via self-introduction detection\n  │\n  └─ [Phase 3: LLM cleanup] ──► transcript.md\n      ├─ LLM speaker role verification (if --speaker-context provided)\n      ├─ Remove fillers (um, uh, 嗯, 啊, etc.)\n      ├─ Fix ASR errors (homophones, context-based)\n      ├─ Polish grammar while preserving meaning\n      └─ (Optional) Identify merged speakers via context\n```\n\n## Language Presets & Models\n\n### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |\n| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |\n| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |\n| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |\n\nSeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)\nto bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.\n\n### `--lang zh-basic` (Chinese, no hotword)\n\nSame as `zh` but uses the base Paraformer-large without hotword support.\nUse when hotword biasing is unnecessary or causing issues with English terms.\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |\n\n### `--lang en` (English)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |\n\n### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |\n\nSenseVoiceSmall includes built-in punctuation and supports emotion detection.\n\n### `--lang whisper` (Multilingual, 99 languages)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |\n\nAll presets share the same VAD (`fsmn-vad`) and speaker diarization (`cam++`) models.\nModels are auto-downloaded from ModelScope on first run.\n\n## Hotword Biasing\n\nSeACo-Paraformer (`--lang zh`) supports hotword customization to improve recognition\nof specific terms — particularly useful for participant names, project names, and\ndomain-specific jargon in meetings.\n\n### How to provide hotwords\n\n```bash\n# Space-separated string\npython3 transcribe_funasr.py meeting.wav --lang zh --hotwords \"张三 李四 ClawCon Rebase\"\n\n# Text file (one word per line)\npython3 transcribe_funasr.py meeting.wav --lang zh --hotwords hotwords.txt\n```\n\n### What to include in hotwords\n\nFor meeting transcription, a good hotwords file includes:\n- **Participant names** (full names in the meeting's language)\n- **Project / product names** (internal codenames, product brands)\n- **Domain-specific Chinese terms** (technical jargon, acronyms in Chinese)\n- **Organization names** (company, team, department names)\n\n### Effectiveness — empirical results\n\nTested on a 4h14m, 9-speaker Chinese meeting with 27 hotwords:\n\n| Term | Without hotwords | With hotwords | Change |\n|------|-----------------|---------------|--------|\n| 龙虾 (lobechat) | 28 | 42 | **+50%** |\n| 高琦 (person name) | 0 | 7 | **0 → 7** |\n| 搬瓦工 (BandwagonHost) | 0 | 1 | **0 → 1** |\n| 谢锐 (person name) | 0 | 1 | improved |\n| 鲲鹏 (org name) | 6 | 7 | slight improvement |\n| Rebase (English) | 5 | 0 | **regression** |\n| Tailwind (English) | 3 | 1 | **regression** |\n\n**Key findings:**\n1. **Chinese terms benefit significantly** — names, brands, and Chinese jargon\n   see clear improvement (龙虾 +50%, 高琦 from zero)\n2. **English terms may regress** — SeACo's hotword biasing operates on Chinese\n   token vocabulary; English loanwords can be disrupted\n3. **Person names have limited uplift** — meeting participants rarely say full\n   names aloud; hotwords help only when names do appear in speech\n\n**Recommendation:** Include Chinese terms and names in hotwords. For English\ntechnical terms, rely on Phase 3 LLM cleanup rather than hotword biasing.\nIf English term accuracy is critical and hotword biasing causes regressions,\nuse `--lang zh-basic` instead.\n\n## Performance Benchmarks\n\nTested on a 4h14m, 9-speaker Chinese meeting recording (GPU: L40S 46GB):\n\n| Metric | Paraformer-large | SeACo + Hotword | Notes |\n|--------|-----------------|-----------------|-------|\n| Model load | 14s | 14s (422s first run, downloading 944MB model) | Cached after first run |\n| Transcription | 169s | 168s | Virtually identical |\n| Raw sentences | 6672 | 6725 | Comparable |\n| Merged segments | 1695 | 1724 | Comparable |\n| Speakers detected | 7 (of 9) | 7 (of 9) | Same diarization result |\n\n| Metric | GPU (L40S 46GB) | CPU (estimated) |\n|--------|-----------------|-----------------|\n| Model load | 14s | ~30s |\n| Transcription | 169s | ~30-60 min |\n| Speaker clustering | ~10s (patched) | ~2-5 min (patched) |\n| LLM cleanup (17 chunks) | ~35 min | ~35 min (network-bound) |\n| Total | ~38 min | ~70-100 min |\n\n**Without the clustering patch**, the original `scipy.linalg.eigh()` on the full Laplacian\nmatrix was O(N^3) and took **10+ hours** on this recording. The patch reduces it to O(N^2*k)\nvia `scipy.sparse.linalg.eigsh()`.\n\n## Clustering Patch (Critical for Long Meetings)\n\nFunASR's `SpectralCluster.get_spec_embs()` uses `scipy.linalg.eigh(L)` which computes\nALL eigenvalues of the NxN Laplacian. For a 4-hour recording, N can be 6000+, making\nthis O(N^3) operation take hours.\n\nThe patch (`scripts/patch_clustering.py`) replaces this with:\n- `scipy.sparse.linalg.eigsh(L_sparse, k=num_speakers, which='SM')` — only computes\n  the k smallest eigenvalues needed, reducing complexity to O(N^2 * k)\n- Vectorized `p_pruning()` — replaces Python loop with numpy broadcasting\n\n**Always run the patch before processing meetings longer than ~1 hour.**\n\n## Speaker Role Verification\n\nSpeaker names are assigned by first-appearance order in the audio, which may\nswap host/guest labels (especially in podcasts). Two layers of automatic\nverification, plus a standalone post-hoc tool:\n\n1. **Phase 2 — self-introduction detection**: scans the first 5 minutes for\n   explicit self-introductions (\"我是X\", \"I'm X\") and swaps labels if mismatched\n2. **Phase 3 — LLM role verification**: when `--speaker-context` is provided,\n   the LLM analyzes the first chunk (up to 15 minutes) before cleanup begins.\n   For 2 speakers: binary CORRECT/SWAP detection. For 3+ speakers: full\n   JSON-based reassignment matching each label to the correct person.\n3. **Post-hoc — `verify_speakers.py`**: standalone script that verifies any\n   existing `*_raw_transcript.json`. Same two modes (2-speaker swap, N-speaker\n   reassignment) with dry-run support. See SKILL.md § Verify Speaker Labels.\n\nFor podcasts, always provide `--speaker-context` describing host/guest roles.\n\n## Speaker Diarization Limitations\n\nFunASR's CAM++ speaker diarization may merge acoustically similar speakers into one ID.\nIn the tested 9-person meeting, only 7 unique IDs were detected (two pairs merged).\n\nWorkarounds:\n1. **Provide `--num-speakers N`** to hint expected count (uses `preset_spk_num`)\n2. **Post-hoc keyword matching**: use reference documents (meeting agendas, attendee notes)\n   to identify which speaker ID maps to which person\n3. **LLM-assisted splitting**: provide `--speaker-context` with per-person keywords;\n   the LLM can then split merged speakers when context is clear (~73% success rate)\n\n## Supporting Files for Better Results\n\nPrepare these files before transcription for best results:\n\n| File | Used in | Purpose |\n|------|---------|---------|\n| `hotwords.txt` | Phase 1 (`--hotwords`) | Bias ASR toward names and terms |\n| `speaker-context.json` | Phase 3 (`--speaker-context`) | Help LLM identify and split speakers |\n| Meeting agenda | Manual reference | Identify meeting phases for post-analysis |\n| Attendee list | Build hotwords + speaker names | Map speaker IDs to real names |\n\n### Example: preparing supporting files from a meeting invite\n\n```bash\n# 1. Create hotwords.txt from attendee list and agenda\ncat > hotwords.txt << 'EOF'\nAlice\nBob\nCarol\nProjectAlpha\nSprint Review\nQ2 OKR\nEOF\n\n# 2. Create speaker-context.json from attendee roles\ncat > speaker-context.json << 'EOF'\n{\n  \"Alice\": \"Engineering manager, discusses sprint velocity and tech debt\",\n  \"Bob\": \"Product manager, presents roadmap and customer feedback\",\n  \"Carol\": \"Designer, shows mockups, mentions Figma and user testing\"\n}\nEOF\n\n# 3. Run with both\npython3 transcribe_funasr.py meeting.wav \\\n  --lang zh --num-speakers 3 \\\n  --speakers \"Alice,Bob,Carol\" \\\n  --hotwords hotwords.txt \\\n  --speaker-context speaker-context.json\n```\n\n## Audio Preprocessing\n\nFunASR works best with 16kHz mono audio. **FLAC is recommended** over WAV — lossless\nquality at ~50% the file size, and FunASR reads it natively via soundfile.\n\n```bash\n# Recommended: FLAC (lossless, compact)\nffmpeg -i recording.m4a -ar 16000 -ac 1 -sample_fmt s16 meeting.flac\n\n# Alternative: WAV (lossless, larger)\nffmpeg -i recording.m4a -ar 16000 -ac 1 meeting.wav\n```\n\n**Important:** Use `-sample_fmt s16` when converting to FLAC — without it, ffmpeg\nmay output 24-bit samples (s32/24bit) which doubles the file size with no ASR benefit.\n\n### Format comparison (4h14m meeting)\n\n| Format | Size | Quality | FunASR support |\n|--------|------|---------|---------------|\n| **FLAC (16kHz mono s16)** | **219MB** | Lossless | Native (soundfile) |\n| WAV (16kHz mono) | 465MB | Lossless | Native (soundfile) |\n| Opus (32kbps) | 54MB | Lossy | Native (soundfile) |\n| M4A/AAC (original 48kHz) | 173MB | Source | Via librosa |\n| M4A/AAC (16kHz 32kbps) | 60MB | Lossy | Via librosa |\n\nFunASR accepts all common audio formats. FLAC offers the best trade-off: lossless\nquality, reasonable size, and native reader support without librosa fallback.\n\nFor long recordings, do NOT split the audio — FunASR handles arbitrarily long files\nand splitting breaks speaker consistency across segments.\n\n## Resume / Checkpoint Support\n\nThe pipeline supports resuming interrupted runs:\n- **Phase 1 output**: `<stem>_raw_transcript.json` — use `--skip-transcribe` to skip ASR\n- **Phase 3 cache**: `<stem>_llm_cache/chunk_NNN.txt` — already-cleaned chunks are reused\n  (kept by default; add `--clean-cache` to delete after completion)\n\n## Model Caching\n\nFunASR models (~3 GB for the `zh` preset) are downloaded from ModelScope on first run\nand cached in `~/.cache/modelscope/hub/`. On ephemeral instances (EC2, cloud VMs), the\ncache is lost when the instance is replaced, requiring a ~2 minute re-download.\n\nTo persist the cache on durable storage (e.g., an EBS data volume):\n\n```bash\n# Via CLI flag (recommended)\npython3 transcribe_funasr.py meeting.flac --model-cache-dir /data/modelscope-cache ...\n\n# Via environment variable\nMODELSCOPE_CACHE=/data/modelscope-cache python3 transcribe_funasr.py meeting.flac ...\n```\n\nThe `systemd-run` examples below include `-E MODELSCOPE_CACHE=...` for this reason.\n\n## Speaker Context JSON Format\n\nThe `--speaker-context` file helps the LLM identify speakers and fix ASR errors:\n\n```json\n{\n  \"Alice\": \"Discussed Q1 revenue targets, mentioned Chicago office relocation\",\n  \"Bob\": \"Presented the new CI/CD pipeline, uses Terraform and ArgoCD\",\n  \"Carol\": \"HR updates, mentioned hiring freeze and new PTO policy\"\n}\n```\n\nThe context is injected into the LLM system prompt for each cleanup chunk.\n\n## Running on CPU-only / Low-Memory Machines\n\nLong recordings (2+ hours) on resource-constrained machines (CPU-only, ≤8 GB RAM)\nface two common failure modes. Both are silent — the process is killed mid-run\nwith no output files saved.\n\n### Problem 1: Process killed by execution timeout\n\nAI coding agents (Claude Code, OpenClaw, Cursor, etc.) impose execution timeouts\non shell commands — typically 2–10 minutes. On a 4-hour recording, CPU transcription\ntakes 1.5–2 hours, well past any agent timeout. The process is silently killed.\n\n**Fix — detach the ASR phase from the agent's process supervision.** Use `--skip-llm`\nfor the detached run because Phase 1 (ASR) is the CPU-intensive bottleneck; Phase 3\n(LLM cleanup) is network-bound and fast — run it afterward via `--skip-transcribe`\nunder the normal agent session.\n\nOption A: `systemd-run` (preferred on systemd hosts):\n\n```bash\nsystemd-run --user --unit=transcribe-job \\\n  --working-directory=/tmp \\\n  -E MODELSCOPE_CACHE=/data/modelscope-cache \\\n  bash -c 'source /path/to/.venv/bin/activate && \\\n    python3 /path/to/transcribe_funasr.py /tmp/meeting.flac \\\n    --lang zh --num-speakers 9 --skip-llm > /tmp/transcribe.log 2>&1'\n\n# Monitor progress\nsystemctl --user status transcribe-job.service\ntail -f /tmp/transcribe.log\n\n# Check result\nls -lh /tmp/*-transcript.md /tmp/*_raw_transcript.json\n```\n\n`systemd-run` creates a transient systemd service fully independent of the agent\nsession — it survives session resets, context pruning, and exec timeouts.\n\n> **Warning:** `systemd-run` creates an isolated mount namespace. FUSE mounts\n> (rclone, sshfs, Google Drive, etc.) from the parent session are NOT visible\n> to the transient service. Copy all dependency files (audio, hotwords,\n> speaker-context, reference documents) to a local path (e.g., `/tmp`) before\n> launching. Always use `--working-directory=/tmp` or another real filesystem path.\n\n> **Note:** `systemd-run --user` requires a user-level systemd instance. It may not\n> work in Docker containers or cloud VMs without `loginctl enable-linger`.\n\nOption B: `nohup` (works everywhere):\n\n```bash\nnohup bash -c 'source .venv/bin/activate && python3 transcribe_funasr.py meeting.flac \\\n  --lang zh --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\necho $!  # Save PID for monitoring\ntail -f transcribe.log\n```\n\n### Problem 2: OOM kill on machines with ≤8 GB RAM\n\nThe `zh` preset loads 4 model components simultaneously (SeACo-Paraformer + VAD +\nPunctuation + CAM++ speaker). Peak RSS can exceed 7 GB on a 4-hour recording. On\nmachines without swap, the OOM killer terminates the process silently.\n\n**Fix A — add swap before running** (requires root):\n\n```bash\nsudo fallocate -l 4G /swapfile\nsudo chmod 600 /swapfile\nsudo mkswap /swapfile\nsudo swapon /swapfile\n\n# After transcription, optionally remove swap\nsudo swapoff /swapfile && sudo rm /swapfile\n```\n\n**Fix B — use `zh-basic` instead of `zh`:**\n\n`zh-basic` (Paraformer-large) loads one fewer model component than `zh`\n(SeACo-Paraformer), reducing peak RSS by ~1–1.5 GB. Accuracy is slightly lower\n(no hotword biasing) but sufficient for most meetings:\n\n```bash\npython3 transcribe_funasr.py meeting.flac --lang zh-basic --num-speakers 9 --skip-llm\n```\n\n**Combining both fixes** (swap + `zh-basic`) reliably handles 4+ hour recordings on\nmachines with as little as 8 GB RAM + 4 GB swap.\n\n### Recommended CPU workflow for long recordings\n\n```bash\n# 1. Add swap if RAM ≤ 8 GB\nsudo fallocate -l 4G /swapfile && sudo chmod 600 /swapfile \\\n  && sudo mkswap /swapfile && sudo swapon /swapfile\n\n# 2. Launch transcription detached from agent timeout\nnohup bash -c 'source .venv/bin/activate && python3 transcribe_funasr.py meeting.flac \\\n  --lang zh-basic --num-speakers 9 --skip-llm' > transcribe.log 2>&1 &\n\n# 3. Monitor\ntail -f transcribe.log\n\n# 4. When done, resume with LLM cleanup (network-bound, runs fine under agent)\npython3 transcribe_funasr.py meeting.flac --skip-transcribe\n```\n\n## Podcast Transcription\n\nThe pipeline handles podcasts and interviews with the same engine, but the workflow\ndiffers from meetings:\n\n### Key differences from meetings\n\n| Aspect | Meeting | Podcast / Interview |\n|--------|---------|---------------------|\n| Speakers | 3–15+, often unknown | 2–3, usually known (host + guests) |\n| Language | Usually single | May mix languages (bilingual hosts) |\n| Hotwords | Participant names + terms | Show name, guest name, topic terms |\n| Speaker context | Role-based keywords | Host asks questions, guest answers |\n| Diarization | Critical | Easier (fewer, distinct voices) |\n\n### Recommended settings\n\n```bash\n# English podcast (2 speakers, host + guest)\npython3 transcribe_funasr.py episode.flac --lang en --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n\n# Bilingual podcast (auto-detect language switches)\npython3 transcribe_funasr.py episode.flac --lang auto --num-speakers 2 \\\n  --speakers \"Alice,Bob\"\n\n# Chinese podcast with topic hotwords\npython3 transcribe_funasr.py episode.flac --lang zh --num-speakers 3 \\\n  --speakers \"主持人,嘉宾A,嘉宾B\" \\\n  --hotwords \"播客名 嘉宾全名 讨论主题关键词\"\n\n# Multi-language podcast (e.g., Spanish + English)\npython3 transcribe_funasr.py episode.flac --lang whisper --num-speakers 2 \\\n  --speakers \"Host,Guest\"\n```\n\n### Tips for podcast transcription\n\n1. **Always provide `--num-speakers`** — podcasts have a known, fixed speaker count;\n   this dramatically improves diarization accuracy with only 2–3 voices\n2. **Always provide `--speakers`** — host/guest names are known upfront\n3. **`--lang auto`** works for bilingual transcript-only output (no speaker labels) —\n   SenseVoiceSmall handles intra-utterance language switching (zh/en/ja/ko/yue) but\n   does not output timestamps, so **speaker diarization is not supported**.\n   Use `--lang zh` for Chinese podcasts that need speaker identification.\n4. **`--lang whisper`** for any other language or heavy code-switching (also lacks\n   timestamp support for diarization — transcript-only)\n5. **Hotwords** — for Chinese podcasts, include the show name and guest's full name;\n   for English podcasts, hotwords are usually unnecessary\n6. **`--speaker-context`** — describe the host/guest dynamic:\n   ```json\n   {\n     \"Alice\": \"Host, asks questions, introduces topics, wraps up segments\",\n     \"Bob\": \"Guest, expert on topic X, shares personal anecdotes\"\n   }\n   ```\n7. **Audio quality** — podcasts are typically studio-recorded with better SNR than\n   meetings; diarization accuracy is correspondingly higher\n8. **Audio source matters** — mobile app downloads may be truncated (trial/preview\n   versions). Download from the web interface for complete files. The script's\n   Phase 0 duration validation catches conversion truncation but cannot detect\n   a source file that is already incomplete.\n\nArchive v1.5.0: 9 files, 51682 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (3814b), scripts/test_speaker_verification.py (59744b), scripts/transcribe_funasr.py (56977b), scripts/verify_speakers.py (17584b), SKILL.md (12079b), _meta.json (150b)\n\nFile v1.5.0:SKILL.md\n\n---\nname: funasr-transcribe\nversion: 1.5.0\ndescription: >\n  This skill should be used when the user explicitly asks to \"transcribe a meeting\",\n  \"transcribe audio\", \"transcribe a meeting recording\",\n  \"convert audio to text\", \"generate meeting minutes from audio\",\n  \"do speech-to-text\", \"transcribe with speaker diarization\",\n  \"identify speakers in audio\", \"transcribe Chinese audio\",\n  \"transcribe English audio\", \"transcribe Japanese audio\",\n  \"multi-speaker transcription\", \"transcribe a podcast\",\n  \"transcribe podcast episode\", \"transcribe an interview\",\n  \"convert podcast to text\", \"podcast to transcript\",\n  or mentions FunASR, Paraformer, SenseVoice, Whisper, meeting\n  transcription, podcast transcription, or speaker diarization.\n  Supports multi-speaker meeting and podcast transcription in Chinese,\n  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper)\n  with automatic speaker diarization and hotword biasing.\n  Works on both GPU and CPU.\nmetadata:\n  openclaw:\n    requires:\n      bins: [\"python3\", \"ffmpeg\"]\n    env_vars:\n      - name: AWS_REGION\n        required: false\n        description: \"AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed.\"\n      - name: ANTHROPIC_API_KEY\n        required: false\n        description: \"API key for Anthropic Claude LLM cleanup\"\n      - name: OPENAI_API_KEY\n        required: false\n        description: \"API key for OpenAI-compatible LLM cleanup\"\n      - name: OPENAI_BASE_URL\n        required: false\n        description: \"Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)\"\n    emoji: \"🎙️\"\n    homepage: \"https://github.com/zxkane/audio-transcriber-funasr\"\n---\n\n# FunASR Meeting & Podcast Transcription\n\nTranscribe multi-speaker audio into structured Markdown with automatic\nspeaker diarization, hotword biasing, and optional LLM cleanup.\n\nAll scripts run directly from the plugin directory — no copying needed.\nDefine this shorthand at the start of every session:\n\n```bash\nSCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/funasr-transcribe/scripts\n```\n\n## Supported Languages\n\n| `--lang` | Model | Languages | Hotword |\n|----------|-------|-----------|---------|\n| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |\n| `zh-basic` | Paraformer-large | Chinese | No |\n| `en` | Paraformer-en | English | No |\n| `auto` | SenseVoiceSmall | Auto-detect: z\n\nArchive v1.4.1: 9 files, 49924 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (3814b), scripts/test_speaker_verification.py (57548b), scripts/transcribe_funasr.py (54001b), scripts/verify_speakers.py (17584b), SKILL.md (11171b), _meta.json (150b)\n\nArchive v1.4.0: 9 files, 46692 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (3814b), scripts/test_speaker_verification.py (46793b), scripts/transcribe_funasr.py (50421b), scripts/verify_speakers.py (17584b), SKILL.md (10992b), _meta.json (150b)\n\nArchive v1.3.1: 9 files, 46642 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (5282b), scripts/setup_env.sh (3814b), scripts/test_speaker_verification.py (46793b), scripts/transcribe_funasr.py (50422b), scripts/verify_speakers.py (17584b), SKILL.md (10689b), _meta.json (150b)\n\nArchive v1.3.0: 9 files, 45237 bytes\n\nFiles: references/pipeline-details.md (19412b), scripts/llm_utils.py (4822b), scripts/patch_clustering.py (3728b), scripts/setup_env.sh (2558b), scripts/test_speaker_verification.py (46793b), scripts/transcribe_funasr.py (50422b), scripts/verify_speakers.py (17584b), SKILL.md (9693b), _meta.json (150b)","readmeExcerpt":"Skill: Audio Transcribe Owner: zxkane Summary: This skill should be used when the user explicitly asks to \"transcribe a meeting\", \"transcribe audio\", \"transcribe a meeting recording\", \"convert audio to te... Tags: latest:1.7.1 Version history: v1.7.1 | 2026-05-01T15:39:21.869Z | auto **MiMo-V2.5-ASR local GPU transcription and major script rename** - Added support for Xiaomi MiMo-V2.5-ASR (8B, local GPU) with improve","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"SCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/audio-transcribe/scripts"},{"language":"bash","snippet":"AUTO_YES=1 bash $SCRIPTS/setup_env.sh\n# Or force CPU:  AUTO_YES=1 bash $SCRIPTS/setup_env.sh cpu"},{"language":"bash","snippet":"# Chinese meeting with hotwords (local-only, no LLM)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt\n\n# English meeting with speaker names\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang en --speakers \"Alice,Bob,Carol,Dave\"\n\n# Auto-detect language (zh/en/ja/ko/yue)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang auto --num-speakers 6\n\n# Whisper for any language\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang whisper --num-speakers 4\n\n# Enable LLM cleanup for polished output (requires --model)\n# Bedrock (uses AWS credential chain: IAM role, SSO, ~/.aws/credentials)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --lang zh --num-speakers 9 --hotwords hotwords.txt \\\n    --provider bedrock --model us.anthropic.claude-sonnet-4-6\n\n# Bedrock \"global\" cross-region profile (recent AWS deployments)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider bedrock --model global.anthropic.claude-sonnet-4-6\n\n# Bedrock via litellm-style wrapper (supported; prefix is stripped for boto3)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider bedrock --model amazon-bedrock/global.anthropic.claude-sonnet-4-6\n\n# Anthropic API (requires ANTHROPIC_API_KEY env var)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider anthropic --model claude-sonnet-4-6\n\n# OpenAI-compatible API (requires OPENAI_API_KEY env var)\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --provider openai --model gpt-4o\n\n# Full pipeline with all supporting files + LLM (best quality)\npython3 $SCRIPTS/transcribe.py episode.m4a \\\n    --lang zh --num-speakers 2 \\\n    --hotwords hotwords.txt \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json \\\n    --reference show-notes.md \\\n    --model us.anthropic.claude-sonnet-4-6\n\n# Resume interrupted LLM cleanup\npython3 $SCRIPTS/transcribe.py meeting.wav \\\n    --skip-transcribe --model us.anthropic.claude-sonnet-4-6"},{"language":"bash","snippet":"# Dry-run: check if host/guest are swapped\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json\n\n# Apply the fix\npython3 $SCRIPTS/verify_speakers.py podcast_raw_transcript.json \\\n    --speakers \"关羽,张飞\" \\\n    --speaker-context speaker-context.json --fix\n\n# Multi-speaker meeting: full reassignment\npython3 $SCRIPTS/verify_speakers.py meeting_raw_transcript.json \\\n    --speakers \"Alice,Bob,Carol,Dave\" \\\n    --speaker-context speaker-context.json --fix\n\n# Then regenerate the markdown with corrected labels\npython3 $SCRIPTS/transcribe.py original.m4a \\\n    --skip-transcribe --clean-cache"},{"language":"bash","snippet":"# One-time: install MiMo on top of the standard environment\nAUTO_YES=1 INSTALL_MIMO=1 \\\n    MIMO_WEIGHTS_PATH=/mnt/models/hf \\\n    bash $SCRIPTS/setup_env.sh"},{"language":"bash","snippet":"python3 $SCRIPTS/transcribe.py podcast.m4a \\\n    --lang mimo --num-speakers 2 \\\n    --mimo-weights-path /mnt/models/hf"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: audio-transcribe\nversion: 1.7.1\ndescription: >\n  This skill should be used when the user explicitly asks to \"transcribe a meeting\",\n  \"transcribe audio\", \"transcribe a meeting recording\",\n  \"convert audio to text\", \"generate meeting minutes from audio\",\n  \"do speech-to-text\", \"transcribe with speaker diarization\",\n  \"identify speakers in audio\", \"transcribe Chinese audio\",\n  \"transcribe English audio\", \"transcribe Japanese audio\",\n  \"multi-speaker transcription\", \"transcribe a podcast\",\n  \"transcribe podcast episode\", \"transcribe an interview\",\n  \"convert podcast to text\", \"podcast to transcript\",\n  or mentions FunASR, Paraformer, SenseVoice, Whisper, MiMo, MiMo-V2.5-ASR,\n  meeting transcription, podcast transcription, or speaker diarization.\n  Supports multi-speaker meeting and podcast transcription in Chinese,\n  English, Japanese, Korean, Cantonese, and 99 languages (via Whisper),\n  plus Xiaomi MiMo-V2.5-ASR (8B, local GPU) for stronger proper-noun and\n  code-switching accuracy. Automatic speaker diarization via CAM++,\n  hotword biasing (FunASR path), LLM cleanup. FunASR works on GPU and CPU;\n  MiMo requires a local CUDA GPU with >=20GB VRAM.\nmetadata:\n  openclaw:\n    requires:\n      bins: [\"python3\", \"ffmpeg\"]\n    env_vars:\n      - name: AWS_REGION\n        required: false\n        description: \"AWS region for Bedrock LLM cleanup (default: us-west-2). Bedrock uses the standard AWS credential chain (IAM role, SSO, ~/.aws/credentials, env vars) — no explicit keys needed.\"\n      - name: ANTHROPIC_API_KEY\n        required: false\n        description: \"API key for Anthropic Claude LLM cleanup\"\n      - name: OPENAI_API_KEY\n        required: false\n        description: \"API key for OpenAI-compatible LLM cleanup\"\n      - name: OPENAI_BASE_URL\n        required: false\n        description: \"Base URL for OpenAI-compatible API (vLLM, Ollama, etc.)\"\n    emoji: \"🎙️\"\n    homepage: \"https://github.com/zxkane/audio-transcriber\"\n---\n\n# Meeting & Podcast Transcription (FunASR + MiMo)\n\nTranscribe multi-speaker audio into structured Markdown with automatic\nspeaker diarization, hotword biasing, and optional LLM cleanup. Two\nASR engine families are available: **FunASR** (Paraformer / SenseVoice /\nWhisper — fast, cheap, GPU or CPU, 99 languages) and **MiMo-V2.5-ASR**\n(Xiaomi's 8B model, local GPU only, stronger on proper nouns and\ncode-switching). Both share the same VAD + speaker-clustering stack.\n\nAll scripts run directly from the plugin directory — no copying needed.\nDefine this shorthand at the start of every session:\n\n```bash\nSCRIPTS=${CLAUDE_PLUGIN_ROOT}/skills/audio-transcribe/scripts\n```\n\n## Supported Languages\n\n| `--lang` | Model | Languages | Hotword |\n|----------|-------|-----------|---------|\n| `zh` (default) | SeACo-Paraformer | Chinese (CER 1.95%) | Yes |\n| `zh-basic` | Paraformer-large | Chinese | No |\n| `en` | Paraformer-en | English | No |\n| `auto` | SenseVoiceSmall | Auto-detect: zh/en/ja/ko/yue | No |\n| `whisper` | Whisper-large-v3-turbo | 99"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn72agp4n0v3y1gk4qds89r0wn82msn3\",\n  \"slug\": \"zxkane-audio-transcriber-funasr\",\n  \"version\": \"1.7.1\",\n  \"publishedAt\": 1777649961869\n}"},{"path":"references/pipeline-details.md","content":"# FunASR Meeting Transcription Pipeline — Technical Details\n\n## Architecture\n\n```\nAudio File (.m4a/.mp3/.wav)\n  │\n  ├─ [ffmpeg] ──► 16kHz mono WAV\n  │\n  ├─ [Phase 1: FunASR] ──► raw_transcript.json\n  │   ├─ FSMN-VAD: segment speech vs silence\n  │   ├─ ASR model (language-dependent, see below)\n  │   ├─ (Optional) Hotword biasing (SeACo-Paraformer only)\n  │   ├─ Punctuation restoration (model-dependent)\n  │   └─ CAM++: speaker embeddings → spectral clustering\n  │\n  ├─ [Phase 2: Post-process]\n  │   ├─ Merge consecutive same-speaker utterances (<2s gap)\n  │   ├─ Map speaker IDs to names (if provided)\n  │   └─ Auto-verify via self-introduction detection\n  │\n  └─ [Phase 3: LLM cleanup] ──► transcript.md\n      ├─ LLM speaker role verification (if --speaker-context provided)\n      ├─ Remove fillers (um, uh, 嗯, 啊, etc.)\n      ├─ Fix ASR errors (homophones, context-based)\n      ├─ Polish grammar while preserving meaning\n      └─ (Optional) Identify merged speakers via context\n```\n\n## Language Presets & Models\n\n### `--lang zh` (Chinese, default) — SeACo-Paraformer with hotword support\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_seaco_paraformer_large_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%), hotword-customizable |\n| VAD | `iic/speech_fsmn_vad_zh-cn-16k-common-pytorch` | 0.4M | Voice activity detection |\n| Punctuation | `iic/punc_ct-transformer_zh-cn-common-vocab272727-pytorch` | 290M | Punctuation restoration |\n| Speaker | `iic/speech_campplus_sv_zh-cn_16k-common` | 7.2M | Speaker diarization |\n\nSeACo-Paraformer accepts a `--hotwords` parameter (space-separated string or .txt file)\nto bias recognition toward specific terms. See [Hotword Biasing](#hotword-biasing) below.\n\n### `--lang zh-basic` (Chinese, no hotword)\n\nSame as `zh` but uses the base Paraformer-large without hotword support.\nUse when hotword biasing is unnecessary or causing issues with English terms.\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-zh-cn-16k-common-vocab8404-pytorch` | 220M | Chinese ASR (CER 1.95%) |\n\n### `--lang en` (English)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/speech_paraformer-large-vad-punc_asr_nat-en-16k-common-vocab10020` | 220M | English ASR |\n\n### `--lang auto` (Auto-detect: zh/en/ja/ko/yue)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/SenseVoiceSmall` | 234M | Multi-language ASR with auto language detection |\n\nSenseVoiceSmall includes built-in punctuation and supports emotion detection.\n\n### `--lang whisper` (Multilingual, 99 languages)\n\n| Component | Model ID | Params | Role |\n|-----------|----------|--------|------|\n| ASR | `iic/Whisper-large-v3-turbo` | 809M | OpenAI Whisper via FunASR, broadest language coverage |\n\nAll presets share the same VAD (`fsmn-vad`) and speaker diari"},{"path":"skill-card.md","content":"## Description:\n\nAudio Transcribe helps an agent transcribe meeting, podcast, and interview recordings into structured transcripts with ASR, speaker diarization, optional speaker verification, and optional LLM cleanup.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zxkane](https://clawhub.ai/user/zxkane)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, operators, and external users can use this skill to convert audio recordings into Markdown transcripts with speaker labels, timestamps, and optional cleanup for meeting notes or podcast/interview review.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Sensitive audio, transcript excerpts, speaker context, or reference material may be sent to external LLM providers when cleanup is enabled.\n\nMitigation: Omit --model or use --skip-llm to keep cleanup local, and only provide trusted reference and speaker-context files when LLM cleanup is needed.\n\nRisk: Speaker-gender inference is enabled by default and may create unnecessary demographic processing.\n\nMitigation: Use --no-detect-gender when demographic inference is not required.\n\nRisk: OpenAI-compatible endpoints and setup commands can change the operational trust boundary.\n\nMitigation: Review any OPENAI_BASE_URL value and run setup commands in an isolated virtual environment after reviewing privileged and dependency behavior.\n\n## Reference(s):\n\n- [Project homepage](https://github.com/zxkane/audio-transcriber)\n- [Pipeline Details](references/pipeline-details.md)\n- [Xiaomi MiMo-V2.5-ASR Model Card](https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR)\n- [ClawHub Skill Page](https://clawhub.ai/zxkane/skills/zxkane-audio-transcriber-funasr)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, JSON, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands; generated transcript output may include Markdown and raw JSON files.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Transcription output can include speaker labels, timestamps, hotword-biased ASR results, and optional LLM-cleaned text.]\n\n## Skill Version(s):\n\n1.7.1 (source: SKILL.md frontmatter and server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1859,"uniquenessScore":41,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T01:27:41.989Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T01:27:41.989Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T03:57:30.131Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}