agentCLAWHUBUnverified

MiniMax Multimodal Toolkit

Generate and process speech, music, video, and images using MiniMax AI with voice cloning, custom voices, multi-scene video, and FFmpeg-based media tools. Skill: MiniMax Multimodal Toolkit Owner: yhlorra Summary: Generate and process speech, music, video, and images using MiniMax AI with voice cloning, custom voices, multi-scene video, and FFmpeg-based media tools. Tags: latest:1.0.0 Version history: v1.0.0 | 2026-03-25T13:43:31.665Z | user Initial publish Archive index: Archive v1.0.0: 18 files, 64004 bytes Files: references/image-api.md (3754b), references/music-api.

OpenClaw

Rank

62

Safety

84

Downloads

2.2k

Updated

Oct 9, 2026

Version

1.0.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.2K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.2K downloadsadoption · observed Oct 9, 2026
Latest release
1.0.0release · observed Mar 25, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s1722rfrg89fk95t1dwm4vqdyx83j73f:yh-minimax-multimodal-toolkit
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-yhlorra-yh-minimax-multimodal-toolkit/snapshot"

Documentation

CLAWHUB

84,897 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: minimax-multimodal-toolkit
description: MiniMax multimodal model skill — use MiniMax  Multi-Modal models for speech, music, video, and image. Create voice, music, video, and images with MiniMax AI: TTS (text-to-speech, voice cloning, voice design, multi-segment), music (songs, instrumentals), video (text-to-video, image-to-video, start-end frame, subject reference, templates, long-form multi-scene), image (text-to-image, image-to-image with character reference), and media processing (convert, concat, trim, extract). Use when the user mentions MiniMax, multimodal generation, or wants speech/music/video/image AI, MiniMax APIs, or FFmpeg workflows alongside MiniMax outputs.
---

# MiniMax Multi-Modal Toolkit

Generate voice, music, video, and image content via MiniMax APIs — the unified entry for **MiniMax multimodal** use cases (audio + music + video + image). Includes voice cloning & voice design for custom voices, image generation with character reference, and FFmpeg-based media tools for audio/video format conversion, concatenation, trimming, and extraction.

## Output Directory

**All generated files MUST be saved to `minimax-output/` under the AGENT'S current working directory (NOT the skill directory).** Every script call MUST include an explicit `--output` / `-o` argument pointing to this location. Never omit the output argument or rely on script defaults.

**Rules:**
1. Before running any script, ensure `minimax-output/` exists in the agent's working directory (create if needed: `mkdir -p minimax-output`)
2. Always use absolute or relative paths from the agent's working directory: `--output minimax-output/video.mp4`
3. **Never** `cd` into the skill directory to run scripts — run from the agent's working directory using the full script path
4. Intermediate/temp files (segment audio, video segments, extracted frames) are automatically placed in `minimax-output/tmp/`. They can be cleaned up when no longer needed: `rm -rf minimax-output/tmp`

## Prerequisites

```bash
brew install ffmpeg jq              # macOS (or apt install ffmpeg jq on Linux)
bash scripts/check_environment.sh
```

No Python or pip required — all scripts are pure bash using `curl`, `ffmpeg`, `jq`, and `xxd`.

### API Host Configuration

MiniMax provides two service endpoints for different regions. Set `MINIMAX_API_HOST` before running any script:

| Region | Platform URL | API Host Value |
|--------|-------------|----------------|
| China Mainland(中国大陆) | https://platform.minimaxi.com | `https://api.minimaxi.com` |
| Global(全球) | https://platform.minimax.io | `https://api.minimax.io` |

```bash
# China Mainland
export MINIMAX_API_HOST="https://api.minimaxi.com"

# or Global
export MINIMAX_API_HOST="https://api.minimax.io"
```

**IMPORTANT — When API Host is missing:**
Before running any script, check if `MINIMAX_API_HOST` is set in the environment. If it is NOT configured:
1. Ask the user which service endpoint their M

_meta.json

{
  "ownerId": "kn7ay31t8b2nm3fjytaw0f9hqh8001cc",
  "slug": "yh-minimax-multimodal-toolkit",
  "version": "1.0.0",
  "publishedAt": 1774446211665
}

references/image-api.md

# MiniMax Image Generation API (image-01)

Source: https://platform.minimaxi.com/docs/api-reference/image-generation-t2i and https://platform.minimaxi.com/docs/api-reference/image-generation-i2i

## Endpoint

`POST https://api.minimaxi.com/v1/image_generation`

## Auth

`Authorization: Bearer <MINIMAX_API_KEY>`

## Request (JSON)

Required:
- `model`: string — `image-01`
- `prompt`: string (max 1500 chars) — text description of the desired image

Optional:
- `aspect_ratio`: string — image aspect ratio, default `1:1`. Options:
  - `1:1` (1024×1024)
  - `16:9` (1280×720)
  - `4:3` (1152×864)
  - `3:2` (1248×832)
  - `2:3` (832×1248)
  - `3:4` (864×1152)
  - `9:16` (720×1280)
  - `21:9` (1344×576)
- `width`: integer — custom width in pixels. Range [512, 2048], must be multiple of 8. Overridden by `aspect_ratio` if both set.
- `height`: integer — custom height in pixels. Same rules as `width`. Both `width` and `height` must be set together.
- `response_format`: string — `url` (default, valid 24h) or `base64`
- `n`: integer (1–9, default 1) — number of images to generate
- `seed`: integer — random seed for reproducibility
- `prompt_optimizer`: boolean (default `false`) — enable automatic prompt optimization
- `aigc_watermark`: boolean (default `false`) — add AIGC watermark

### Subject Reference (image-to-image)

- `subject_reference`: array — character reference for image-to-image generation
  - `type`: string — currently only `character` (portrait)
  - `image_file`: string — reference image as public URL or Base64 Data URL (`data:image/jpeg;base64,...`). For best results, use a single person front-facing photo. Formats: JPG, JPEG, PNG. Max size: 10MB.

## Example — Text-to-Image

```json
{
  "model": "image-01",
  "prompt": "A man in a white t-shirt, full-body, standing front view, outdoors, with the Venice Beach sign in the background, Los Angeles. Fashion photography in 90s documentary style, film grain, photorealistic.",
  "aspect_ratio": "16:9",
  "response_format": "url",
  "n": 3,
  "prompt_optimizer": true
}
```

## Example — Image-to-Image (Character Reference)

```json
{
  "model": "image-01",
  "prompt": "A girl looking into the distance from a library window",
  "aspect_ratio": "16:9",
  "subject_reference": [
    {
      "type": "character",
      "image_file": "https://example.com/face.jpg"
    }
  ],
  "n": 2
}
```

## Response

```json
{
  "id": "03ff3cd0820949eb8a410056b5f21d38",
  "data": {
    "image_urls": ["https://...", "https://...", "https://..."],
    "image_base64": null
  },
  "metadata": {
    "success_count": 3,
    "failed_count": 0
  },
  "base_resp": {
    "status_code": 0,
    "status_msg": "success"
  }
}
```

- `data.image_urls`: array of image URLs (when `response_format` is `url`, valid 24h)
- `data.image_base64`: array of Base64 strings (when `response_format` is `base64`)
- `metadata.success_count`: number of successful

references/music-api.md

# MiniMax Music Generation API (music-2.5)

Source: https://platform.minimaxi.com/docs/api-reference/music-generation

## Endpoint

`POST https://api.minimaxi.com/v1/music_generation`

## Auth

`Authorization: Bearer <MINIMAX_API_KEY>`

## Request (JSON)

Required:
- `model`: string — `music-2.5`
- `lyrics`: string (1–3500 chars) — required. Use `\n` for line breaks. Structure tags: `[Verse]`, `[Chorus]`, `[Bridge]`, `[Intro]`, `[Outro]`, etc.

Optional:
- `prompt`: string (0–2000 chars) — style description, optional but recommended.
- `lyrics_optimizer`: boolean — auto-generate lyrics from prompt when lyrics is empty.
- `stream`: boolean (default `false`)
- `output_format`: `hex` (default) or `url`. URL valid for 24 hours.
- `aigc_watermark`: boolean — top-level field, non-streaming only.
- `audio_setting`:
  - `sample_rate`: 16000, 24000, 32000, 44100
  - `bitrate`: 32000, 64000, 128000, 256000
  - `format`: mp3, wav, pcm

## Example

```json
{
  "model": "music-2.5",
  "prompt": "indie folk, melancholic, introspective",
  "lyrics": "[verse]\n...\n[chorus]\n...",
  "aigc_watermark": false,
  "audio_setting": {
    "sample_rate": 44100,
    "bitrate": 256000,
    "format": "mp3"
  }
}
```

## Response

- `data.audio`: hex string or URL depending on `output_format`
- `data.status`: 1 (generating), 2 (complete)
- `extra_info`: duration, sample_rate, channels, bitrate, size
- `base_resp.status_code`: 0 on success

## Notes

- `music-2.5` does not support `is_instrumental`. For instrumental music, use lyrics `[intro] [outro]` and add `pure music, no lyrics` to the prompt.
- `prompt` is optional but recommended for better style control.
- `stream=true` only supports `hex` output.

references/tts-guide.md

# TTS Guide

## Setup

```bash
cd skills/MiniMaxStudio
pip install -r requirements.txt
brew install ffmpeg   # macOS (or: sudo apt install ffmpeg)
export MINIMAX_API_KEY="your-api-key"   # sk-api-xxx or sk-cp-xxx
python scripts/check_environment.py
```

## Quick Test

```bash
python scripts/tts/generate_voice.py tts "Hello, this is a test." -o test.mp3
```

## Voice Management

List available voices:

```bash
python scripts/tts/generate_voice.py list-voices
```

### Voice Cloning

Create a custom voice from an audio sample:

```bash
python scripts/tts/generate_voice.py clone audio.mp3 --voice-id my-custom-voice

# With preview
python scripts/tts/generate_voice.py clone audio.mp3 --voice-id my-voice --preview "Test text" --preview-output preview.mp3
```

Requirements: 10s–5min duration, ≤20MB, mp3/wav/m4a format.

### Voice Design

Design a voice from a text description:

```bash
python scripts/tts/generate_voice.py design "A warm, gentle female voice" --voice-id designed-voice
```

Custom voices expire after 7 days if not used with TTS.

## Audio Processing

### Merge

```bash
python scripts/tts/generate_voice.py merge file1.mp3 file2.mp3 -o combined.mp3
python scripts/tts/generate_voice.py merge a.mp3 b.mp3 -o merged.mp3 --crossfade 300
```

### Convert

```bash
python scripts/tts/generate_voice.py convert input.wav -o output.mp3
python scripts/tts/generate_voice.py convert input.wav -o output.mp3 --format mp3 --bitrate 192k --sample-rate 32000
```

FFmpeg required. Supported formats: mp3, wav, flac, ogg, m4a, aac, wma, opus, pcm.

## Segment-Based TTS

For multi-voice, multi-emotion workflows using a `segments.json` file:

```bash
# Validate
python scripts/tts/generate_voice.py validate segments.json --verbose

# Generate
python scripts/tts/generate_voice.py generate segments.json -o output.mp3 --crossfade 200
```

### segments.json Format

```json
[
  { "text": "Hello!", "voice_id": "female-shaonv", "emotion": "" },
  { "text": "How are you?", "voice_id": "male-qn-qingse", "emotion": "happy" }
]
```

- `text` (required): Text to synthesize
- `voice_id` (required): Voice ID
- `emotion` (optional): For speech-2.8 models, leave empty for auto-matching. Valid values: happy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisper

## Troubleshooting

| Error | Solution |
|-------|----------|
| `MINIMAX_API_KEY is required` | `export MINIMAX_API_KEY="key"` |
| `FFmpeg not installed` | `brew install ffmpeg` |
| `Voice not found` | `python scripts/tts/generate_voice.py list-voices` |
| `401 Unauthorized` | Check API key validity |
| `429 Too Many Requests` | Add delays between requests |

## API Details

- **Endpoint**: `POST /v1/t2a_v2`
- **Base URL**: `https://api.minimaxi.com`
- **Auth**: `Authorization: Bearer {MINIMAX_API_KEY}`
- **Models**: speech-2.8-hd (recommended), speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd
Github ReposUpdated 5h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/yhlorra/skills/yh-minimax-multimodal-toolkit",
      "sourceUrl": "https://clawhub.ai/yhlorra/skills/yh-minimax-multimodal-toolkit",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:18:50.741Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-yhlorra-yh-minimax-multimodal-toolkit/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-yhlorra-yh-minimax-multimodal-toolkit/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:18:50.741Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.2K downloads",
      "href": "https://clawhub.ai/yhlorra/yh-minimax-multimodal-toolkit",
      "sourceUrl": "https://clawhub.ai/yhlorra/yh-minimax-multimodal-toolkit",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:18:50.741Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.0",
      "href": "https://clawhub.ai/yhlorra/yh-minimax-multimodal-toolkit",
      "sourceUrl": "https://clawhub.ai/yhlorra/yh-minimax-multimodal-toolkit",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-03-25T13:43:31.665Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-yhlorra-yh-minimax-multimodal-toolkit/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-yhlorra-yh-minimax-multimodal-toolkit/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.0",
      "description": "Initial publish",
      "href": "https://clawhub.ai/yhlorra/yh-minimax-multimodal-toolkit",
      "sourceUrl": "https://clawhub.ai/yhlorra/yh-minimax-multimodal-toolkit",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-03-25T13:43:31.665Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to MiniMax Multimodal Toolkit and adjacent AI workflows.