
Video Prompt Director
- 1 installs
- 4 repo stars
- Updated May 13, 2026
- zhouwei713/video-prompt-director
Turns a short idea, brief, link, or product docs into a source-grounded 30-second video generation prompt with scenes, screen text, narration, and constraints.
About
Extracts facts from provided or researched sources and converts them into a controlled, copy-ready video generation prompt for HyperFrames or other video tools. A developer uses it when turning a rough video idea into a detailed shot-by-shot prompt with duration, screen text, motion, and negative constraints.
- Produces 4-6 scenes for a 30-second video with exact screen text and narration
- Grounds every claim in provided or researched sources and labels creative assumptions
Video Prompt Director by the numbers
- 1 all-time installs (skills.sh)
- Ranked #1,200 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
npx skills add https://github.com/zhouwei713/video-prompt-director --skill video-prompt-directorAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 4 |
| Last updated | May 13, 2026 |
| Repository | zhouwei713/video-prompt-director ↗ |
What it does
Turns a short idea, brief, link, or product docs into a source-grounded 30-second video generation prompt with scenes, screen text, narration, and constraints.
Files
Video Prompt Director
Purpose
Turn a short idea into a high-control video generation prompt. The output must be ready to copy into HyperFrames or another video generation workflow, with clear duration, format, scenes, screen text, visual actions, motion, narration, and source-grounded details.
Default video length is 30 seconds unless the user requests another duration or the platform clearly implies a different length.
Core Workflow
1. Identify the request type. Product demo, tool introduction, teaching short, social short, historical explainer, concept explainer, brand intro, event recap, or pitch video.
2. Gather source material. Use user-provided text first. If the user gives links, asks to research, or asks for current facts, inspect sources before writing. Prefer official websites, product docs, GitHub repositories, release notes, English documentation, academic or primary sources. Avoid Chinese websites unless the user supplied them or explicitly requested them.
3. Extract video-worthy facts. Pull the product name, core promise, top functions, commands, UI labels, workflow steps, outputs, audience, differentiators, numbers, limitations, and visual assets. See references/source-extraction.md when sources are involved.
4. Convert facts into scenes. Every important fact must become at least one visible item: screen text, typed input, UI chip, card, timeline marker, graph, icon, comparison row, cursor action, or narration line.
5. Write the controlled prompt. Use the output contract below. Keep it in Chinese when the user writes in Chinese. Include only facts that are grounded in provided or researched material. If a useful detail is inferred, label it as a creative assumption.
6. Check the prompt. Confirm the prompt has a clear goal, 4 to 6 scenes for 30 seconds, exact screen text, motion instructions, style constraints, audio guidance, and negative constraints.
Output Contract
Return one copy-ready video prompt with these sections:
视频类型
Choose the best type and name it plainly.
时长与画幅
Default: 30 秒,16:9 横屏。 Adjust only if the user asks or the distribution channel implies it.
核心目标
One sentence: what the viewer should understand, remember, or do after watching.
资料提取摘要
List the facts that will be used in the video. Include commands, feature names, UI labels, data points, and outputs. If no source was available, list assumptions under 创意假设.
叙事主线
One sentence describing the flow from opening to closing.
分镜提示词
Write 4 to 6 scenes. Each scene should include:
1. 时间段 2. 画面 3. 屏幕文字 4. 动作与镜头 5. 旁白 6. 转场
关键屏幕文字
List every exact phrase that must appear on screen. Preserve commands and product terms exactly. For command-like text, mark tags and regular input separately.
动效与镜头控制
Describe UI actions, typing, clicks, generation state, zooms, transitions, camera movement, progress, and object motion.
风格要求
Set visual tone, surface design, realism level, color direction, typography, and density.
禁止项
Include constraints that prevent drift: no unsupported claims, no overdone effects, no clutter, no tiny text, no fake UI if a real UI is required, no arbitrary feature invention.
最终可复制提示词
Provide a single consolidated prompt that can be pasted into a video generation agent.
Defaults
1. Duration: 30 seconds. 2. Scenes: 5 scenes for most topics. 3. Opening: Show the central object or result within the first 3 seconds. 4. Ending: Repeat the product, topic, or key takeaway with one concise line. 5. Style: Clean, premium, controlled motion. For product and UI videos, use realistic interface behavior. For history and education, use timelines, maps, documents, diagrams, and labels. 6. Audio: Include concise narration and optional captions. Captions should match the narration.
Source Handling
Use references/source-extraction.md when the prompt depends on provided material, links, documents, or web research.
Use references/prompt-patterns.md when selecting scene patterns or adapting the example prompt structure.
Quality Bar
A good output should make the video feel directed, not merely described. It should say what appears, what moves, what text is typed or displayed, what the viewer hears, and what the model must avoid. The prompt must preserve grounded details such as /鲸格skills 制作ppt when they are central to the story.
interface:
display_name: "Video Prompt Director"
short_description: "Turn briefs and sources into controlled video prompts"
default_prompt: "Use $video-prompt-director to turn this one-line idea into a source-grounded 30-second video generation prompt."
policy:
allow_implicit_invocation: true
Video Prompt Director
把一句视频想法、产品简介、链接资料或研究主题,转成更有控制力的视频生成提示词。
This Codex skill turns a short video idea, product brief, source link, or research topic into a controlled video generation prompt.
中文说明
适合场景
1. 产品功能演示 2. 工具介绍 3. 教学短片 4. 社媒短视频 5. 历史或商业解读 6. 基于资料的视频脚本扩写
核心能力
1. 默认生成 30 秒视频提示词。 2. 支持从用户给定资料中提取重点信息。 3. 支持在需要时自行查找资料,并优先使用官方文档、英文资料、GitHub 仓库、论文、新闻源和其他权威来源。 4. 会把重点功能、命令、产品卖点、数据点、界面文字和视觉钩子转成视频文案。 5. 输出包含分镜、屏幕文字、旁白、镜头动作、动效控制、风格要求和禁止项。 6. 适合继续交给 HyperFrames 或其他视频生成工具执行。
使用示例
$video-prompt-director 介绍一个能自动生成 PPT 和视频的工具$video-prompt-director 根据这个链接,为一个 GitHub 开源项目制作 30 秒介绍视频提示词:https://github.com/heygen-com/hyperframes$video-prompt-director 介绍一下豆包 APP 如果收费,对于国内大模型市场的影响输出内容
1. 视频类型 2. 时长与画幅 3. 核心目标 4. 资料提取摘要 5. 叙事主线 6. 分镜提示词 7. 关键屏幕文字 8. 动效与镜头控制 9. 风格要求 10. 禁止项 11. 最终可复制提示词
资料处理原则
1. 优先使用用户提供的资料。 2. 如果用户提供链接或要求研究,会先提取来源中的关键事实。 3. 对产品类内容,重点提取产品名称、核心功能、用户动作、输出结果、差异化卖点和界面文字。 4. 对历史、商业和市场内容,重点提取时间、主体、事件、数据、因果关系和可视化线索。 5. 对不确定内容,会标记为创意假设,避免写成确定事实。
English
What It Does
Video Prompt Director converts a brief idea into a structured, source grounded video generation prompt. It is designed for Codex workflows that need a stronger prompt before creating videos with HyperFrames or another video tool.
Best For
1. Product demos 2. Tool introductions 3. Teaching shorts 4. Social videos 5. Historical explainers 6. Market analysis videos 7. Source based video scripts
Main Features
1. Uses 30 seconds as the default duration. 2. Extracts key facts from user provided materials. 3. Researches sources when needed, with preference for official docs, English language sources, GitHub repositories, papers, reputable news, and primary sources. 4. Converts functions, commands, product benefits, data points, UI labels, and visual hooks into video copy. 5. Produces scene by scene prompts with narration, screen text, motion direction, camera control, style rules, and negative constraints. 6. Works well as a prompt preparation step before HyperFrames video production.
Example Prompts
$video-prompt-director Create a 30 second prompt for a tool that can automatically generate slides and videos.$video-prompt-director Turn this GitHub repo into a product intro video prompt: https://github.com/heygen-com/hyperframes$video-prompt-director Explain how paid access for a leading AI app could affect the China large model market.Output Sections
1. Video type 2. Duration and aspect ratio 3. Core objective 4. Source extraction summary 5. Narrative through line 6. Scene prompts 7. Required screen text 8. Motion and camera control 9. Style requirements 10. Negative constraints 11. Final copy ready prompt
File Structure
video-prompt-director/
SKILL.md
README.md
agents/
openai.yaml
references/
prompt-patterns.md
source-extraction.mdInstallation
Copy this folder into your Codex skills directory:
$CODEX_HOME/skills/video-prompt-directorThen invoke it in Codex with:
$video-prompt-director your video idea hereDesign Notes
This skill focuses on direction quality. A good output should clearly define what appears on screen, what moves, what text is shown, what the viewer hears, what facts are grounded in sources, and what the video generator must avoid.
Prompt Patterns
Use these patterns to turn extracted information into a controlled video generation prompt.
Product Demo Pattern
Best for tools, plugins, apps, SaaS features, agents, and workflows.
Scene structure for 30 seconds:
1. Result hook, 0 to 4 seconds Show the final output or multiple outputs immediately.
2. Real interface setup, 4 to 9 seconds Show the product surface or realistic app screen.
3. User action, 9 to 15 seconds Type a command, upload a file, click a button, select a mode, or invoke an agent.
4. Generation or processing, 15 to 21 seconds Show status, progress, icon motion, logs, cards, or previews.
5. Output inspection, 21 to 27 seconds Zoom into the generated artifact, open a preview, or compare before and after.
6. Closing memory line, 27 to 30 seconds Show product name and one sharp takeaway.
Control details to include:
1. Exact typed text. 2. Tag-like treatment for commands or plugin names. 3. Cursor movement and click targets. 4. Loading or generation motion. 5. Output preview and camera push-in. 6. Clean transition style.
Teaching Short Pattern
Best for explaining a concept, process, technical idea, or method.
Scene structure:
1. Problem or question. 2. Core concept. 3. Step-by-step breakdown. 4. Example in action. 5. Summary and takeaway.
Use diagrams, labels, callouts, progress rails, and simple examples.
Historical Explainer Pattern
Best for events, biographies, and social movements.
Scene structure:
1. Context and pressure. 2. Trigger event. 3. Turning point. 4. Escalation or conflict. 5. Resolution. 6. Legacy.
Use dates, maps, documents, portraits only when sourced, symbolic objects, and restrained motion.
Social Short Pattern
Best for short vertical hooks.
Scene structure:
1. A strong first-frame question or result. 2. Three quick proof beats. 3. One action line.
Keep text large, scenes fast, and captions synchronized.
Example Pattern From User
The provided example works because it defines:
1. Video type. 2. Duration and format. 3. Core objective. 4. Scene-by-scene UI progression. 5. Exact screen text. 6. Exact typed command. 7. Visual distinction between plugin tag and normal input. 8. Generation state. 9. Output preview and camera push-in. 10. Style and negative constraints.
Reusable abstraction:
1. Start with proof of capability. 2. Move to a real-looking product surface. 3. Show one precise user input. 4. Make the product process visible. 5. Show the generated output. 6. End with product name and memory line.
Screen Text Rules
1. Put exact phrases in quotes. 2. Preserve commands exactly. 3. Separate tag text from user text when UI identity matters. 4. Avoid long paragraphs on screen. 5. Repeat the product or topic name near the end.
Motion Rules
1. Use motion to explain state change. 2. Prefer typing, clicking, sliding, zooming, expanding, folding, progress, and preview opening. 3. Use smooth fades for UI transitions. 4. Avoid excessive particle effects or abstract motion when the video is about a product workflow. 5. Every scene should have one dominant motion and one secondary ambient motion.
Source Extraction
Use this reference when the user provides materials, links, files, product descriptions, or asks the agent to find information before writing a video prompt.
Priority Order
1. User-provided material. 2. Official product website, docs, help center, GitHub repo, changelog, release notes. 3. Primary or authoritative sources for non-product topics. 4. English-language reputable sources for background context.
Avoid Chinese websites unless the user directly provides them or explicitly asks to use them.
Extract These Items
1. Name Product, topic, feature, person, event, or organization.
2. One-line definition What it is in plain language.
3. Core functions or claims The top 3 to 6 capabilities that the video should show.
4. User action Commands, prompts, clicks, uploads, selections, typed input, API calls, or setup steps.
5. Output What the tool or process creates: PPT, video, report, chart, dashboard, lesson, plan, artifact.
6. Differentiator Why this matters compared with a generic workflow.
7. Evidence Numbers, dates, versions, status, license, supported platforms, source quotes, or documented behavior.
8. Visual hooks UI surfaces, icons, documents, maps, cards, timelines, screenshots, graphs, before-and-after states.
9. Screen text Exact words that must appear in the video.
10. Constraints What cannot be claimed, shown, or implied.
Convert Information Into Video Content
For each extracted item, decide how it appears:
1. Product function becomes a typed command, feature chip, UI card, generated artifact, or narration sentence. 2. A workflow step becomes cursor movement, click, upload, progress state, or timeline marker. 3. A number becomes a large stat, badge, counter, or comparison row. 4. A document becomes a paper panel, page preview, highlighted sentence, or source card. 5. A historical event becomes a date marker, map, document, symbol, or motion path.
Handling Weak Sources
If a source is incomplete, do not invent specifics. Use 创意假设 for plausible staging choices and keep factual claims conservative.
If a claim is uncertain, phrase it as a direction for the video, not as a fact.
If the user wants a real product interface, preserve real UI labels and avoid imaginary controls unless they are clearly marked as stylized.
Minimum Source Summary
Before writing the final prompt, create a compact extraction summary:
1. 主题 2. 目标观众 3. 关键事实 4. 必须出现的屏幕文字 5. 可视化素材 6. 不应声称的内容
This summary should feed the final video prompt directly.