English Learning Animation
Produce and quality-check animated English shorts with Qwen3-TTS.
Rank
62
Safety
84
Downloads
1.1k
Updated
Oct 11, 2026
Version
0.1.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 0.1.0release · observed Aug 26, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s174rz3f0862tcfw7pzfh5w8kn83hv2z:english-learning-animation- Install using `clawhub skill install s174rz3f0862tcfw7pzfh5w8kn83hv2z:english-learning-animation` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/tobewin/english-learning-animation before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-english-learning-animation/snapshot"
Documentation
CLAWHUB
28,166 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: english-learning-animation description: Create or revise short, character-led English-learning animation videos with an English-only hook/cover, original editorial-cartoon visuals, distinct Qwen3-TTS VoiceDesign characters, and audio-driven scene timing. Use when the user asks for animated English lessons, dialogue-based language-learning shorts, cartoon ESL videos, or to improve their character voices, subtitles, cover, motion, or audiovisual synchronization. --- # English Learning Animation Create a coherent, original short-form English lesson. Prioritize a watchable scene over a slide deck with narration. ## Workflow 1. Define one communicative outcome and write an English-only script with 3–5 usable phrases, a natural dialogue, and a brief repeat-after-me close. Do not choose a target runtime first. Let the generated speech, necessary pauses, cover, and recap determine the final duration. Many lessons will naturally land near 25–45 seconds, but this is not a quota. 2. Build a shot list before generating visuals. Add a `semantic_contract` to `script.json`: topic, setting, visual brief, required scene tags, and stale terms that must never appear. Every scene needs `semantic_tags`. Give every voice segment a stable semantic owner such as `customer`, `barista`, or `narrator`; do not use gender as the long-term character identity. 3. Use Qwen3-TTS VoiceDesign when no reference audio exists. Create a short audition for every recurring role first; do not reuse one voice for multiple characters. Lock each approved role's `voice_profile` and add line-specific `performance` direction. 4. Generate an empty background plate and separate transparent character/prop cutouts that visibly match the current setting. A hotel lobby cannot stand in for a subway station, restaurant, or attraction. Use a layered animation system such as `paper-collage-remotion`; never animate a single flattened illustration as the whole video. 5. Place audio using actual generated durations, then make visual changes at segment starts, phrase beats, and turn changes. Never add dead air or extend scenes merely to reach a round-number runtime. During dialogue, keep the speaker visually primary; during narration, use an English phrase card or semantic graphic rather than pretending a character is speaking. 6. Render only after passing the quality gates below. ## Production Starter Initialize a new project from the approved layered-animation baseline: ```bash python <skill>/scripts/init_project.py <new-or-empty-project-directory> ``` Add the empty background plates and transparent cutouts at the asset paths declared in `script.json`; do not copy generated user content into the skill. Each speaking cutout should declare a `speaker` field matching the narration role id. Generate role-separated audio with the bundled script: ```bash python <skill>/scripts/generate_qwen3_voices.py voice-manifest.json \ --model <local-qwen3-tts-voicedesign-checkpoint> ``` Use `voice-ma
README.md
# English Learning Animation
Create short, character-led English-learning videos with layered editorial-cartoon visuals, role-matched Qwen3-TTS voices, English-only on-video copy, and audio-driven timing.
It is designed for repeatable social-video production: a clear cover, a small practical dialogue, phrase practice, and an export that has passed both mechanical checks and visual review.
## What it produces
- Original paper-collage / editorial-cartoon scenes composed from a background plate and independent transparent character layers.
- Distinct Qwen3 VoiceDesign roles, each with a stable voice profile and per-line performance cue.
- Actual-audio-driven scene timing rather than an arbitrary target runtime.
- An English-only video surface, including cover, captions, and practice cards.
- A preflight and post-render review pipeline for sync, layer opacity, render streams, and review-frame extraction.
## Real output
Hotel Wi-Fi and breakfast lesson cover:

Representative review frames from the same episode:

Breakfast-order episode review frames:

## Quality guarantees
The skill validates the production constraints that commonly break short animated lessons:
- Role ownership, voice-profile stability, audio duration, segment windows, and no overlap.
- A solid-alpha character matte: no ghosted characters or flattened background plates.
- English-only on-video text by default and a 2–3 second outcome-led cover.
- Low-frequency, low-amplitude character motion to avoid visual shaking.
- Topic-to-scene metadata via `semantic_contract` and `semantic_tags`.
- Data-driven phrase cards from `script.json`, preventing stale cards from earlier episodes.
Visual meaning still needs human review. The validation workflow extracts a cover and one representative frame for every spoken segment so that setting, props, speaker emphasis, phrase cards, and captions can be checked before publishing.
## Quick start
```bash
python scripts/init_project.py /path/to/new-lesson
cd /path/to/new-lesson
```
Edit `voice-manifest.json` and `script.json`, add the required background plate and transparent character layers, then generate role-separated audio:
```bash
python /path/to/english-learning-animation/scripts/generate_qwen3_voices.py \
voice-manifest.json \
--model /path/to/Qwen3-TTS-VoiceDesign
```
Run the complete preflight before rendering:
```bash
python /path/to/english-learning-animation/scripts/validate_project.py .
```
After the Remotion render:
```bash
python /path/to/english-learning-animation/scripts/validate_project.py . \
--video out/final.mp4 \
--review-dir work/review-frames
```
## Required lesson metadata
Each project declares its topic and visual intent in `script.json`:
```json
{
"semantic_contrac_meta.json
{
"ownerId": "kn75z6gevjsyrznm7dg2ez6sen82h8sz",
"slug": "english-learning-animation",
"version": "0.1.0",
"publishedAt": 1787724032742
}references/quality-gates.md
# Quality Gates ## Voice - One stable Qwen3 VoiceDesign instruction per character. - Use semantic role ids (`customer`, `barista`) and keep gender/age inside the voice profile rather than using them as identity keys. - Character age, role, and delivery are visibly distinct in an audition. - Every line combines a stable `voice_profile` with a line-specific `performance` cue. - Audio duration is read from the generated file before timeline placement. ## Sync - Voice manifest filenames, speakers, text, cover title, and cover duration exactly match the Remotion timeline. - Each audio segment has a visual owner and a matching start frame. - The narration role id matches a declared layer `speaker` or a cutout filename. - Dialogue lines foreground the speaking character; narrator lines foreground an English learning graphic. - No audio begins before its intended picture is visible. - Captions or phrase cards begin and end with their audio segment. - Declared audio windows contain the full measured waveform and do not overlap. ## Motion - Background, characters, and props are separate layers. - All character cutouts have a real alpha channel, transparent exterior pixels, and at least 88% fully opaque visible pixels. - Reject a cutout if more than 12% of its visible pixels are semitransparent; this is the mechanical ghosting alarm. - Speaking motion is subtle and low frequency; no jitter, flicker, or high-frequency scale oscillation. - At least 4 distinct visual beats for a normal dialogue lesson; add more beats when the content needs them. - Final runtime follows the generated speech and intentional pauses. Do not pad, trim, or hold a scene merely to hit a preferred integer or range. ## Semantic continuity - `script.json` declares a `semantic_contract` and each scene carries topic-relevant `semantic_tags`. - Each setting receives a setting-appropriate plate and props; do not repurpose a previous location merely because its visual style matches. - Phrase cards are data-driven from `script.json`. No card may survive from a starter or earlier episode unless it belongs to the current lesson. - Before release, inspect the cover and every extracted speech frame for setting, character role, phrase-card, caption, and dialogue consistency. ## Publishing - Cover is 2–3 seconds, English-only, and states the learning outcome. - Omit runtime from the cover by default. If the user explicitly requests a runtime claim, derive it from the final render and require it to match. - Video surface contains no Chinese unless the user explicitly requests bilingual on-video subtitles. - Validate resolution, codec, audio stream, and total duration with `ffprobe`. - Extract the cover and one midpoint frame for every spoken segment; inspect character solidity, speaker emphasis, subtitle correctness, and composition.
skill-card.md
## Description: Create or revise short, character-led English-learning animation videos with an English-only hook or cover, original editorial-cartoon visuals, distinct Qwen3-TTS VoiceDesign characters, and audio-driven scene timing. This skill is ready for commercial/non-commercial use. ## Publisher: [tobewin](https://clawhub.ai/user/tobewin) ### License/Terms of Use: MIT ## Use Case: Developers and creators use this skill to produce short English-learning animation videos with dialogue, phrase practice, role-separated voices, layered visuals, and preflight and post-render quality checks. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Editable lesson files can direct some scripts to write or inspect files outside the intended project folder. Mitigation: Use the skill only with trusted lesson projects in a sandboxed working directory, keep voice-manifest.json output_dir and segment file values project-relative, avoid absolute paths or '..', and do not run validators on untrusted script.json files. Risk: Generated lessons can pass mechanical checks while still showing the wrong setting, props, speaker emphasis, phrase cards, or captions. Mitigation: Inspect the extracted cover and one representative frame per spoken segment before publishing. ## Reference(s): - [Source repository](https://github.com/ToBeWin/english-learning-animation) - [ClawHub skill page](https://clawhub.ai/tobewin/skills/english-learning-animation) - [Quality Gates](references/quality-gates.md) - [README](README.md) ## Skill Output: **Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with JSON contracts, shell commands, and generated project files] **Output Parameters:** [1D] **Other Properties Related to Output:** [Outputs guide local video production and validation; model weights, generated voices, and final videos are not bundled.] ## Skill Version(s): 0.1.0 (source: server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/tobewin/skills/english-learning-animation",
"sourceUrl": "https://clawhub.ai/tobewin/skills/english-learning-animation",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T12:31:41.467Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-english-learning-animation/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-english-learning-animation/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T12:31:41.467Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.1K downloads",
"href": "https://clawhub.ai/tobewin/english-learning-animation",
"sourceUrl": "https://clawhub.ai/tobewin/english-learning-animation",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T12:31:41.467Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.0",
"href": "https://clawhub.ai/tobewin/english-learning-animation",
"sourceUrl": "https://clawhub.ai/tobewin/english-learning-animation",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-26T06:00:32.742Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-english-learning-animation/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-english-learning-animation/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.0",
"description": "- Initial release of the skill for English-learning animation video creation. - Supports generating short, character-led English lesson animations with unique editorial-cartoon visuals and distinct Qwen3-TTS voices. - Introduces a defined workflow: scripting with communicative outcomes, visual/voice asset management, layered animation, and strict language/visual rules. - Includes quality gate validations and reproducible production tools for script, visual, audio, and final video outputs. - Provides starter scripts and guidance for asset and voice management, ensuring English-only, immersive lessons with editable contracts and review frames.",
"href": "https://clawhub.ai/tobewin/english-learning-animation",
"sourceUrl": "https://clawhub.ai/tobewin/english-learning-animation",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-26T06:00:32.742Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
