Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
bytedance avatar

Video Generation

  • 2.5k installs
  • 79.3k repo stars
  • Updated August 5, 2026
  • bytedance/deer-flow

video-generation creates AIGC clips from JSON prompt files and generate.py with Gemini Veo or MiniMax backends.

About

The video-generation skill creates AIGC videos using structured JSON prompts and a bundled Python generate.py script. Workflow step one gathers subject, style, technical specs, aspect ratio, and optional reference images without scanning user folders under /mnt/user-data. Step two writes a descriptive JSON prompt file to /mnt/user-data/workspace/. Step three optionally creates a reference image via the image-generation skill; a single image guides the first or last frame. Step four runs generate.py with --prompt-file, optional --reference-images, --output-file, and --aspect-ratio defaulting to 16:9. The skill forbids reading the Python source; call it with parameters only. Provider selection is automatic: GEMINI_API_KEY enables Gemini Veo; only MINIMAX_API_KEY routes to MiniMax async video with first_frame_image; VIDEO_GENERATION_PROVIDER can force gemini or minimax. Outputs land in /mnt/user-data/outputs/ and should be shared via present_files with brief result notes and iteration offers. Prompts must be English regardless of user language.

  • Structured JSON prompts written to /mnt/user-data/workspace/ before generation.
  • generate.py called with prompt file, reference images, output path, aspect ratio.
  • Do not read generate.py; invoke with CLI parameters only.
  • GEMINI_API_KEY or MINIMAX_API_KEY auto-selects provider; MiniMax ignores aspect-ratio.
  • Optional reference image from image-generation skill guides first or last frame.

Video Generation by the numbers

  • 2,508 all-time installs (skills.sh)
  • +61 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #141 of 1,335 Generative Media skills by installs in the Skillselion catalog
  • Security screen: CRITICAL risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

video-generation capabilities & compatibility

Capabilities
json prompt authoring for scene, camera, dialogu · generate.py cli execution with reference images · gemini veo and minimax provider auto selection · output delivery via present_files with iteration
Works with
openai
Use cases
video generation · image generation · orchestration
Runs
Runs locally
Pricing
Bring your own API key
From the docs

What video-generation says it does

Do NOT read the python file, instead just call it with the parameters.
SKILL.md
Always use English for prompts regardless of user's language
SKILL.md
npx skills add https://github.com/bytedance/deer-flow --skill video-generation

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs2.5k
repo stars79.3k
Security audit0 / 3 scanners passed
Last updatedAugust 5, 2026
Repositorybytedance/deer-flow

How do I generate a short video clip from a structured scene description with optional reference frames?

Generate AIGC videos from structured JSON prompts and optional reference images via generate.py with Gemini Veo or MiniMax providers.

Who is it for?

Agent sessions that need scripted video clips with camera, dialogue, and audio blocks in JSON.

Skip if: Skip for static image-only requests; use image-generation instead.

When should I use this skill?

User asks to generate, create, or imagine videos with structured prompts or reference images.

What you get

An MP4 in /mnt/user-data/outputs/ produced from a validated JSON prompt and optional reference image.

  • MP4 video file
  • Python integration function

Files

SKILL.mdMarkdownGitHub ↗

Video Generation Skill

Overview

This skill generates high-quality videos using structured prompts and a Python script. The workflow includes creating JSON-formatted prompts and executing video generation with optional reference image.

Core Capabilities

  • Create structured JSON prompts for AIGC video generation
  • Support reference image as guidance or the first/last frame of the video
  • Generate videos through automated Python script execution

Workflow

Step 1: Understand Requirements

When a user requests video generation, identify:

  • Subject/content: What should be in the image
  • Style preferences: Art style, mood, color palette
  • Technical specs: Aspect ratio, composition, lighting
  • Reference image: Any image to guide generation
  • You don't need to check the folder under /mnt/user-data

Step 2: Create Structured Prompt

Generate a structured JSON file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}.json

Step 3: Create Reference Image (Optional when image-generation skill is available)

Generate reference image for the video generation.

  • If only 1 image is provided, use it as the guided frame of the video

Step 3: Execute Generation

Call the Python script:

python /mnt/skills/public/video-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/prompt-file.json \
  --reference-images /path/to/ref1.jpg \
  --output-file /mnt/user-data/outputs/generated-video.mp4 \
  --aspect-ratio 16:9

Parameters:

  • --prompt-file: Absolute path to JSON prompt file (required)
  • --reference-images: Absolute paths to reference image (optional)
  • --output-file: Absolute path to output image file (required)
  • --aspect-ratio: Aspect ratio of the generated image (optional, default: 16:9)

[!NOTE] Do NOT read the python file, instead just call it with the parameters.

Video Generation Example

User request: "Generate a short video clip depicting the opening scene from "The Chronicles of Narnia: The Lion, the Witch and the Wardrobe"

Step 1: Search for the opening scene of "The Chronicles of Narnia: The Lion, the Witch and the Wardrobe" online

Step 2: Create a JSON prompt file with the following content:

{
  "title": "The Chronicles of Narnia - Train Station Farewell",
  "background": {
    "description": "World War II evacuation scene at a crowded London train station. Steam and smoke fill the air as children are being sent to the countryside to escape the Blitz.",
    "era": "1940s wartime Britain",
    "location": "London railway station platform"
  },
  "characters": ["Mrs. Pevensie", "Lucy Pevensie"],
  "camera": {
    "type": "Close-up two-shot",
    "movement": "Static with subtle handheld movement",
    "angle": "Profile view, intimate framing",
    "focus": "Both faces in focus, background soft bokeh"
  },
  "dialogue": [
    {
      "character": "Mrs. Pevensie",
      "text": "You must be brave for me, darling. I'll come for you... I promise."
    },
    {
      "character": "Lucy Pevensie",
      "text": "I will be, mother. I promise."
    }
  ],
  "audio": [
    {
      "type": "Train whistle blows (signaling departure)",
      "volume": 1
    },
    {
      "type": "Strings swell emotionally, then fade",
      "volume": 0.5
    },
    {
      "type": "Ambient sound of the train station",
      "volume": 0.5
    }
  ]
}

Step 3: Use the image-generation skill to generate the reference image

Load the image-generation skill and generate a single reference image narnia-farewell-scene-01.jpg according to the skill.

Step 4: Use the generate.py script to generate the video

python /mnt/skills/public/video-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/narnia-farewell-scene.json \
  --reference-images /mnt/user-data/outputs/narnia-farewell-scene-01.jpg \
  --output-file /mnt/user-data/outputs/narnia-farewell-scene-01.mp4 \
  --aspect-ratio 16:9
Do NOT read the python file, just call it with the parameters.

Output Handling

After generation:

  • Videos are typically saved in /mnt/user-data/outputs/
  • Share generated videos (come first) with user as well as generated image if applicable, using present_files tool
  • Provide brief description of the generation result
  • Offer to iterate if adjustments needed

Notes

  • Always use English for prompts regardless of user's language
  • JSON format ensures structured, parsable prompts
  • Reference image enhance generation quality significantly
  • Iterative refinement is normal for optimal results

Providers (Gemini / MiniMax)

Auto-selected by environment variables (CLI unchanged):

  • GEMINI_API_KEY set → Gemini Veo (default, unchanged).
  • Only MINIMAX_API_KEY set → MiniMax video (/v1/video_generation, async 3-step poll/download).
  • Force with VIDEO_GENERATION_PROVIDER=gemini|minimax.

MiniMax overrides: MINIMAX_API_HOST (default https://api.minimaxi.com), MINIMAX_VIDEO_MODEL (default MiniMax-Hailuo-2.3). The first reference image is used as MiniMax first_frame_image. MiniMax ignores --aspect-ratio (it uses resolution/duration).

Related skills

How it compares

Choose this when you already use Deer Flow and need a minimal Python wrapper around Veo instead of building the REST payload manually.

FAQ

Which API keys select the video provider?

GEMINI_API_KEY selects Gemini Veo; only MINIMAX_API_KEY selects MiniMax. Override with VIDEO_GENERATION_PROVIDER.

Should I read the generate.py source?

No. Call the script with --prompt-file, --reference-images, --output-file, and --aspect-ratio only.

What language should JSON prompts use?

Always English regardless of the user's language.

Is Video Generation safe to install?

skills.sh reports 0 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

Generative Mediaautomationllm

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.