AI短剧制作助手 | AI Short Film Producer
AI短剧制作助手 | AI Short Film Producer — 低成本AI短剧/短片全流程制作技能。使用Grok Imagine生成视频镜头、TTS生成配音,配合FFmpeg+Python本地合成。适用于从零制作AI短片、短视频、短剧EP、预告片等场景。包含完整的分镜脚本创作、视频生成、配音生成、音频驱动... Skill: AI短剧制作助手 | AI Short Film Producer Owner: hitjcl Summary: AI短剧制作助手 | AI Short Film Producer — 低成本AI短剧/短片全流程制作技能。使用Grok Imagine生成视频镜头、TTS生成配音,配合FFmpeg+Python本地合成。适用于从零制作AI短片、短视频、短剧EP、预告片等场景。包含完整的分镜脚本创作、视频生成、配音生成、音频驱动... Tags: ai:1.0.3, film:1.0.3, latest:1.0.3, production:1.0.3, video:1.0.3 Version history: v1.0.3 | 2026-04-30T11:23:17.316Z | user 标题改为中英文;移除外部链接和具体定价避免安全误报 v1.0.2 | 2026-04-30T09:50:26.768Z | use
Rank
62
Safety
84
Downloads
3.6k
Updated
Oct 9, 2026
Version
1.0.3
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 3.6K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 3.6K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.0.3release · observed Apr 30, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17fcraf4fka3zh3rs6gjkpky983gfj1:ai-short-film-producer- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-hitjcl-ai-short-film-producer/snapshot"
Documentation
CLAWHUB
98,849 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: ai-short-film-producer
description: AI短剧制作助手 | AI Short Film Producer — 低成本AI短剧/短片全流程制作技能。使用Grok Imagine生成视频镜头、TTS生成配音,配合FFmpeg+Python本地合成。适用于从零制作AI短片、短视频、短剧EP、预告片等场景。包含完整的分镜脚本创作、视频生成、配音生成、音频驱动剪辑、字幕叠加、最终合成、成本核算的全套SOP。
---
# AI短剧制作助手 | AI Short Film Producer
## 概述
本Skill提供一套完整的**低成本AI短剧制作流程**,从脚本创作到最终成片,总成本仅需**¥30-50/部**(128秒短片)。核心思路:用AI API生成素材 → 本地FFmpeg合成 → WorkBuddy编排调度。
**适用场景:**
- 用户说"帮我做一个短片/短剧/预告片"
- 用户说"把这段文案做成视频"
- 用户说"生成一个XX题材的短视频"
- 用户需要从零到一完成AI视频制作
**核心成本优势:**
- 视频生成:Grok Imagine(速创API,按秒计费)
- 配音生成:TTS(速创API,按字计费)
- 合成剪辑:本地FFmpeg免费
- AI编排:WorkBuddy Lite版
---
## 制作流程总览
```
Step 1: 脚本创作
├── 确定主题/时长/风格
├── 编写分镜脚本(镜头×台词×角色)
└── 输出:分镜表 + TTS文本清单
Step 2: 视频镜头生成
├── 调用速创API Grok Imagine
├── 25个镜头批量异步生成
└── 输出:ep1_shots/*.mp4
Step 3: TTS配音生成
├── 调用速创API audio_tts
├── 多角色多音色
└── 输出:ep1_tts/*.mp3
Step 4: 音频驱动剪辑
├── 逐段按TTS时长裁剪/循环镜头
├── 短镜头自动stream_loop填充
└── 输出:分段seg_*.mp4
Step 5: 字幕生成
├── Python Pillow生成透明PNG字幕
├── FFmpeg overlay叠加(因FFmpeg 8.x无drawtext)
└── 输出:带字幕的分段视频
Step 6: 最终合成
├── concat拼接25段视频
├── concat拼接25段音频
├── 音视频合并
└── 输出:最终成片.mp4
Step 7: 素材导出
├── 结构化桌面文件夹
├── 矩阵表 + JSON
└── 成本核算
```
---
## 详细步骤
### Step 1: 脚本创作
**输入:** 用户需求(主题、风格、时长、参考素材)
**输出:** 分镜脚本文档 + TTS台词清单
**工作流程:**
1. 与用户确认主题方向(科幻/悬疑/科普/剧情等)
2. 编写分镜脚本,包含:
- 镜头编号、画面描述、时长
- 配音台词、角色分配、音色选择
- 音效说明
3. 输出TTS台词清单(25段以内,每段2-20字最佳)
4. 角色音色分配表:
| 角色类型 | 推荐音色ID | 说明 |
|---------|-----------|------|
| 旁白/叙述者 | male-qn-jingying | 精英青年男声,通用 |
| 男主角 | male-qn-jingying | 精英青年男声 |
| 霸道/硬汉 | male-qn-badao | 霸道男声 |
| 反派/俊朗 | junlang_nanyou | 俊朗男声 |
| 成熟女性 | female-chengshu | 成熟女声 |
| 少女 | female-shaonv | 少女音 |
| 研究员/学生 | male-qn-daxuesheng | 大学生男声 |
| 醇厚长辈 | male-chunhou | 醇厚男声 |
### Step 2: 视频镜头生成(速创API Grok Imagine)
**API平台:** 速创API(详见 references/sucuang_api.md)
**模型:** Grok Imagine(xAI Aurora引擎)
**价格:** 按秒计费(具体见平台)
**API调用方式:**
- 鉴权:Authorization Header 传API Key(不带Bearer前缀)
- 接口:POST /api/async/video/grok_imagine
- 参数格式:扁平JSON
- 结果查询:GET /api/async/detail?id=xxx(轮询直到status=2)
**批量生成策略:**
1. 25个镜头同时提交(用ThreadPoolExecutor)
2. 每个镜头约10秒,生成耗时约30-60秒
3. 失败自动重试(平均重试3次)
4. 注意:Sora2接口已不可用(持续400错误),全部使用Grok Imagine
**Prompt编写要点:**
- 英文Prompt效果更稳定
- 包含:场景描述、光线、构图、镜头运动
- 示例:`"Deep space, Milky Way galaxy slowly rotating, cinematic wide shot, photorealistic, 4K quality"`
### Step 3: TTS配音生成(速创API audio_tts)
**API接口:** POST /api/async/audio_tts
**价格:** 按字计费(具体见平台)
**参数格式(重要):** 扁平JSON,不要嵌套
```json
{
"text": "台词内容",
"voice_id": "male-qn-jingying",
"speed": 1.0
}
```
**注意事项(踩坑经验):**
- ❌ 不要传 format 参数(会报500"存在未绑定的参数")
- ❌ 不要嵌套成 `{"model":"audio_tts","params":{...}}`
- ✅ 状态码判断:status=2 完成,status=0/1 处理中
- ⚠️ 部分任务会卡住(status一直=0),重试可换IP节点
- ✅ 返回tar包,需解压获取mp3
### Step 4: 音频驱动剪辑(核心节奏控制)
**核心理念:** 画面长度由语音旁白决定,而非固定时长。先录制/生成TTS配音,再让每段视频精确匹配对应配音的时长。这样保证音画天然同步,且节奏由配音自然驱动。
#### 4.1 节奏控制逻辑
```
每段(镜头, TTS)的处理流程:
1. 获取TTS音频实际时长 tts_dur(用ffprobe精确到毫秒)
2. 获取源视频时长 src_dur
3. 对比决策:
├── src_dur >= tts_dur + 0.5s → 直接裁剪到_meta.json
{
"ownerId": "kn703gwc11aqnnaatk69f339d9822e08",
"slug": "ai-short-film-producer",
"version": "1.0.3",
"publishedAt": 1777548197316
}references/production_workflow.md
# AI短剧制作流程详细参考
## 一、项目启动
### 需求确认
与用户确认以下信息:
1. **主题**:科幻/悬疑/科普/剧情/搞笑/教育等
2. **风格**:写实/动画/赛博朋克/古风等
3. **时长**:30秒/60秒/2分钟/5分钟
4. **角色**:需要几个角色?是否有特定音色要求?
5. **参考素材**:用户是否提供文案/图片/参考视频?
### 脚本模板
```markdown
# 项目名称 · 分镜脚本
## 总览
- 总时长:XX秒
- 镜头数:XX个
- 角色数:XX个
## 分镜表
| # | 镜头ID | 画面描述 | 配音台词 | 角色 | 音色 | 时长 |
|---|--------|---------|---------|------|------|------|
| 1 | shot01_xxx | 画面描述 | 台词内容 | 角色名 | 音色ID | 10s |
| 2 | shot02_xxx | 画面描述 | 台词内容 | 角色名 | 音色ID | 10s |
```
---
## 二、视频生成
### Grok Imagine Prompt编写指南
**结构模板:**
```
[场景描述], [主体描述], [光线/氛围], [构图/镜头运动], [风格关键词]
```
**示例:**
```
"Deep space, Milky Way galaxy slowly rotating across a star-filled cosmos,
countless stars twinkling, nebulae in deep blue and purple hues,
cinematic wide shot, photorealistic, 4K quality, slow gentle rotation"
```
**各类型场景Prompt关键词:**
| 场景类型 | 关键词 |
|---------|--------|
| 太空/宇宙 | deep space, stars, nebula, galaxy, cosmic, celestial |
| 城市/街头 | city street, urban, bustling, modern, buildings |
| 室内/会议室 | conference room, modern, screens, professional |
| 人物特写 | close-up, portrait, expression, emotion, cinematic lighting |
| 自然/风景 | landscape, mountains, ocean, sunset, cinematic |
| 科技/实验室 | laboratory, technology, screens, equipment, futuristic |
### 批量提交策略
1. 25个镜头同时提交到API
2. 使用ThreadPoolExecutor(max_workers=10)
3. 每5秒轮询一次结果
4. 失败自动重试(最多5次)
5. 下载后检查文件完整性(ffprobe验证时长)
---
## 三、TTS配音生成
### 多角色音色分配策略
| 角色数量 | 分配策略 |
|---------|---------|
| 1-2个角色 | 旁白用jingying,对话角色用对应音色 |
| 3-5个角色 | 主要角色各分配独立音色,次要角色复用 |
| 6-10个角色 | 核心角色独立音色,路人/群演用通用音色 |
### 台词字数控制
- 每段TTS建议2-20字(太长影响听感)
- 10秒镜头配2-4秒台词最合适
- 留白时间给观众消化内容
---
## 四、音频驱动剪辑(核心节奏控制)
### 核心理念
**音频驱动剪辑 = 画面长度由语音旁白决定,而非固定时长。**
传统剪辑思维是"先定视频长度,再往里塞配音",结果配音节奏被画面绑架。音频驱动反过来——先录制TTS配音,再让每段视频精确匹配对应配音的时长。这样:
- 观众听到的每句话都有完整的画面时长
- 叙事节奏由台词自然驱动
- 不会出现"话没说完画面就切了"
### 核心算法
```
对于每一段(镜头, TTS):
1. 获取TTS音频时长 tts_dur(ffprobe精确到毫秒)
2. 获取源视频时长 src_dur
3. 对比决策:
├── src_dur >= tts_dur + 0.5s → 直接裁剪到tts_dur
├── src_dur ≈ tts_dur(差<0.5s)→ 直接裁剪
└── src_dur < tts_dur → stream_loop循环填充
4. 输出:seg_NNN.mp4(时长精确=tts_dur)
```
### 逐段精确裁剪(避免累积漂移)
```python
import subprocess
from pathlib import Path
FFPROBE = '/opt/homebrew/bin/ffprobe'
FFMPEG = '/opt/homebrew/bin/ffmpeg'
def get_duration(filepath):
"""获取媒体文件精确时长"""
result = subprocess.run(
[FFPROBE, '-v', 'quiet', '-show_entries', 'format=duration',
'-of', 'csv=p=0', str(filepath)],
capture_output=True, text=True
)
return float(result.stdout.strip())
def process_segment(tts_file, shot_file, output_file):
"""处理单段:按TTS时长裁剪/循环视频"""
tts_dur = get_duration(tts_file)
src_dur = get_duration(shot_file)
cmd = [FFMPEG, '-y']
if src_dur >= tts_dur:
# 视频够长,直接裁剪
cmd.extend(['-t', str(tts_dur), '-i', str(shot_file)])
else:
# 视频不够长,循环填充
cmd.extend(['-stream_loop', '-1', '-i', str(shot_file),
'-t', str(tts_dur)])
cmd.extend(['-c:v', 'libreferences/sucuang_api.md
# 速创API 接口文档与踩坑经验
## 平台信息
- **平台地址**: (注册后获取)
- **文档中心**: (注册后获取)
- **API Key获取**: 注册登录后进入控制台获取
## 通用鉴权方式
**推荐方式 — Authorization Header(不带Bearer前缀):**
```python
HEADERS = {
"Authorization": "你的API_KEY",
"Content-Type": "application/json"
}
```
**❌ 不要用URL参数传key:**
```python
# 会返回403,不要这样用
requests.get("平台API地址/api/xxx?key=你的API_KEY")
```
---
## 1. Grok Imagine 视频生成
### 接口信息
- **接口**: POST `/api/async/video/grok_imagine`
- **价格**: 按秒计费(具体见平台)
- **点数**: 5点/秒
- **免费额度**: 无
- **QPS限制**: 100次/秒
- **每日限制**: 付费用户不限制
### 请求参数
```json
{
"prompt": "英文描述效果更稳定,包含场景、光线、构图、镜头运动",
"duration": 10,
"style": "cinematic"
}
```
### 结果查询
- **接口**: `GET /api/async/detail?id=xxx`
- **轮询策略**: 每5秒查询一次
- **状态码**: status=2 表示完成,status=0/1 表示处理中
- **返回内容**: 包含视频下载URL
### 批量生成策略
```python
from concurrent.futures import ThreadPoolExecutor, as_completed
def submit_shot(shot):
# 提交生成任务
resp = requests.post(url, headers=HEADERS, json=params)
task_id = resp.json()["data"]["id"]
# 轮询直到完成
while True:
result = requests.get(f"{BASE}/detail?id={task_id}", headers=HEADERS)
if result.json()["data"]["status"] == 2:
return download_video(result.json()["data"]["video_url"])
time.sleep(5)
# 批量25个镜头同时提交
with ThreadPoolExecutor(max_workers=10) as executor:
futures = [executor.submit(submit_shot, shot) for shot in SHOTS]
for future in as_completed(futures):
results.append(future.result())
```
### 踩坑经验
- Sora2接口(sora2/video)已不可用,持续返回400错误,全部使用Grok Imagine
- 英文Prompt比中文Prompt效果更稳定
- 平均重试3次才能获得满意结果
- 生成耗时约30-60秒/个
---
## 2. TTS配音生成
### 接口信息
- **接口**: `POST /api/async/audio_tts`
- **价格**: 按字计费(具体见平台)
- **返回格式**: tar包(需解压获取mp3)
### 请求参数(重要:扁平JSON)
```json
{
"text": "台词内容",
"voice_id": "male-qn-jingying",
"speed": 1.0
}
```
### 可用音色列表
| 音色ID | 描述 | 适用角色 |
|--------|------|---------|
| male-qn-jingying | 精英青年男声 | 旁白、汪淼、常伟思 |
| male-qn-badao | 霸道男声 | 史强 |
| male-qn-daxuesheng | 大学生男声 | 研究员 |
| male-chunhou | 醇厚男声 | 长辈角色 |
| junlang_nanyou | 俊朗男声 | 潘寒 |
| female-chengshu | 成熟女声 | 申玉菲 |
| female-shaonv | 少女音 | 杨冬 |
### 踩坑经验(重要)
1. **❌ 不要传 format 参数** — 会报500"存在未绑定的参数"
2. **❌ 不要嵌套参数** — 不要写成 `{"model":"audio_tts","params":{...}}`,直接扁平JSON
3. **✅ 状态码判断** — 用 `status == 2` 判断完成,不要用字符串 `"completed"`
4. **⚠️ 部分任务会卡住** — status一直=0,重试可换到不同IP节点
5. **✅ 返回tar包** — 需要用 `tarfile` 解压获取mp3文件
### TTS生成代码模板
```python
import requests, time, tarfile, io
def generate_tts(text, voice_id, output_path):
payload = {"text": text, "voice_id": voice_id, "speed": 1.0}
resp = requests.post(f"{BASE}/api/async/audio_tts", headers=HEADERS, json=payload)
task_id = resp.json()["data"]["id"]
while True:
result = requests.get(f"{BASE}/api/async/detail?id={task_id}", headers=HEADERS)
data = result.json()["data"]
if data["status"] == 2: # 完成
audio_url = data["audio_url"]
audio_resp = requests.get(audio_url)
tar = tarfile.open(fileobj=io.BytesIO(audio_skill-card.md
## Description: AI短剧制作助手 | AI Short Film Producer helps agents plan and produce low-cost AI short films by creating shot scripts, generating video and TTS assets through third-party APIs, and assembling them locally with FFmpeg and Python. This skill is ready for commercial/non-commercial use. ## Publisher: [hitjcl](https://clawhub.ai/user/hitjcl) ### License/Terms of Use: MIT-0 ## Use Case: External users, developers, and creative operators use this skill to turn a topic, script, or concept into an AI short film workflow with shot planning, video generation, voiceover generation, subtitles, local composition, review, and cost tracking. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The workflow uses third-party video and TTS APIs that may receive user prompts, scripts, or generated content. Mitigation: Do not submit secrets, regulated data, private scripts, or sensitive personal information to the referenced APIs; review provider terms before use. Risk: API keys may be exposed if copied into URLs, logs, scripts, or shared project files. Mitigation: Pass API keys through headers or a local secret store, keep them out of URLs and logs, and avoid committing generated credentials or configuration files. Risk: Downloaded media archives and generated assets are processed locally with Python and FFmpeg. Mitigation: Use request timeouts, retry limits, trusted download hosts, archive size checks, and safe FFmpeg concat manifests before processing untrusted media folders. Risk: Generated footage, TTS, or subtitles can be inaccurate, poorly synchronized, or visually defective. Mitigation: Run the documented review workflow for audio-video sync, subtitle accuracy, duration drift, repeated loops, and visual quality before publishing. ## Reference(s): - [AI Short Film Producer skill page](https://clawhub.ai/hitjcl/skills/ai-short-film-producer) - [AI短剧制作流程详细参考](artifact/references/production_workflow.md) - [速创API 接口文档与踩坑经验](artifact/references/sucuang_api.md) ## Skill Output: **Output Type(s):** [Guidance, Markdown, Code, Shell commands, Configuration] **Output Format:** [Markdown with JSON, Python, and shell command examples] **Output Parameters:** [1D] **Other Properties Related to Output:** [Produces workflow guidance and implementation snippets for scripts, generated media assets, subtitles, FFmpeg composition, review checks, and cost calculations.] ## Skill Version(s): 1.0.3 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/hitjcl/skills/ai-short-film-producer",
"sourceUrl": "https://clawhub.ai/hitjcl/skills/ai-short-film-producer",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T07:37:12.770Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-hitjcl-ai-short-film-producer/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-hitjcl-ai-short-film-producer/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T07:37:12.770Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "3.6K downloads",
"href": "https://clawhub.ai/hitjcl/ai-short-film-producer",
"sourceUrl": "https://clawhub.ai/hitjcl/ai-short-film-producer",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T07:37:12.770Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.3",
"href": "https://clawhub.ai/hitjcl/ai-short-film-producer",
"sourceUrl": "https://clawhub.ai/hitjcl/ai-short-film-producer",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-30T11:23:17.316Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-hitjcl-ai-short-film-producer/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-hitjcl-ai-short-film-producer/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.3",
"description": "标题改为中英文;移除外部链接和具体定价避免安全误报",
"href": "https://clawhub.ai/hitjcl/ai-short-film-producer",
"sourceUrl": "https://clawhub.ai/hitjcl/ai-short-film-producer",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-30T11:23:17.316Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
