
Fpv Immersive Video Prompting
- 9 installs
- 83 repo stars
- Updated June 10, 2026
- zhouwei713/fpv-immersive-video-prompting
Helps with ai & agent building tasks.
About
fpv-immersive-video-prompting is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted coding.
- fpv-immersive-video-prompting
- AI & Agent Building
- AI-coding skill
Fpv Immersive Video Prompting by the numbers
- 9 all-time installs (skills.sh)
- +1 installs in the week ending Jul 28, 2026 (Skillselion tracking)
- Ranked #12,152 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
npx skills add https://github.com/zhouwei713/fpv-immersive-video-prompting --skill fpv-immersive-video-promptingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 9 |
|---|---|
| repo stars | ★ 83 |
| Last updated | June 10, 2026 |
| Repository | zhouwei713/fpv-immersive-video-prompting ↗ |
What it does
Helps with ai & agent building tasks.
Files
FPV Immersive Video Prompting
Overview
This skill turns a scene idea plus optional image references into a directed FPV AI-video prompt. The target output should feel like the viewer is moving through a playable scene, not watching a generic beauty shot.
Use this for image-to-video workflows where the user wants:
- one-shot first-person walkthroughs
- numbered route stops or camera planning
- characters reacting as the camera approaches
- human or non-human POVs: guest, pet cat, drone, robot vacuum, bird, vehicle, object, spirit
- variable numbers of characters or targets
- strong identity and route consistency
Core Principle
Default to Chinese for all generated prompts, including GPT Image asset prompts, route-control image prompts, video prompts, negative prompts, and delivery notes, unless the user explicitly asks for English or a target tool requires English. Keep technical terms such as FPV, Seedance, Kling, Runway, Veo, one-shot, prompt, and path control in English when they are clearer.
Write the output as a small playable scene specification:
1. Who or what is the camera? 2. Where does it start? 3. How many main targets are there exactly? 4. What numbered order does it visit them in? 5. How does the camera physically move? 6. What happens at each stop? 7. What must stay consistent? 8. What must never appear?
Default Asset Workflow
When the user needs GPT Image / GPT-Image-2 assets, request a multi-image asset pack rather than one crowded contact sheet. Treat requests for “生图 prompt”, “首帧图”, “参考图”, “素材包”, “用 GPT Image 生成”, or any image-prep step as an asset-pack request, not as a single first-frame prompt.
For any case that needs multiple images, output one batch-generation prompt that asks GPT Image / ChatGPT to generate multiple separate images in a single response. Do not split the asset pack into many independent per-image prompts unless the user explicitly asks for that. The prompt must say: “请一次性生成 [X] 张独立图片,不要生成拼图、九宫格、contact sheet 或一张图里塞多个画面。”
For close-interaction scenes with N main people/targets, always output the complete asset pack unless the user explicitly asks for only one image: 1. First-frame scene image with small numbered stop markers only: 1, 2, 3, ... N 2. N separate character/reference images, one for each main person/target 3. Optional clean first frame without numbers for safer image-to-video input
Before returning any image prompt, run this count check: total images = 1 route/numbered first frame + N character references + optional 1 clean first frame. The video prompt must refer to these images by role, e.g. “图片 1 is route planning, 图片 2-4 are character references, 图片 5 is clean first frame.”
Default image sets:
Close-interaction / character mode: 1. First-frame scene image with small numbered stop markers only: 1, 2, 3, ... 2. One character/reference image for each main target 3. Optional clean first frame without numbers for safer image-to-video input
World-route / path-control mode: 1. Aerial route-control map or wide first frame with one clear continuous red path from start to destination 2. Optional clean world/scene reference without the red path 3. Optional landmark or character references only if specific destinations need consistent appearance
Prefer numbered stop markers for close-range character interactions, indoor scenes, crowded social scenes, and exact target-count workflows. GPT Image often struggles to draw one continuous physically coherent route through furniture, people, walls, railings, or water; a bad red line can mislead the video model. Numbered stops are easier to generate correctly, and the video prompt can define the movement between them.
Use red-line path control for large-scale route-shaped scenes: aerial world maps, fantasy continent journeys, city-to-landmark flythroughs, racing lines, canyon/drone routes, open-world game trailers, and Seedance 2.0 path-control demos. In that mode, the route image is a planning artifact; the final video must remove all red lines, arrows, annotations, labels, and map-view appearance.
Red-Line Path Control Rules
Use this mode when the user asks for Seedance 2.0 path control, a drawn route, an aerial map flythrough, a game-world traversal, a racing line, or a large-scale drone/bird/spirit FPV journey.
For the route-control image prompt:
- Use a high-resolution 16:9 aerial terrain map, world map, tactical map, or wide route-planning scene.
- Draw one clear continuous red route line, optionally with a subtle arrow, from start to destination.
- Make the route physically plausible through visible corridors: roads, valleys, rivers, city gates, bridges, canyons, rooftops, airspace, tunnels, coastlines, or ridgelines.
- Build 4-6 visually distinct route segments so the video has progression: peaceful start, transition zone, landmark/city, danger zone, final destination.
- Keep the red line clean and legible. Avoid multiple competing paths unless the user explicitly wants branching routes.
For the video prompt:
- Say the uploaded image is a route-planning map/image, not the final look.
- Say the red line/arrow/annotations are only camera-path controls and must be completely removed.
- State the camera must strictly follow the drawn route geometry from start to destination.
- Use invisible first-person drone, bird, spirit, vehicle, or mounted-camera POV unless the route is truly walkable.
- Add timeline segments by geography, not by character stop.
- Include camera language: continuous forward motion, natural banking, close passes, altitude changes, foreground parallax, progressive acceleration, smooth horizon control.
- Avoid map view, visible red lines, annotations, teleporting, reverse motion, jump cuts, visible drone, guide characters, flat terrain, deformed landmarks, flickering structures.
Use numbered stops instead when the video depends on close character interactions, exact character count, indoor navigation, or object-level continuity.
Numbered Stop Marker Rules
For the first-frame image prompt:
- Do not draw red route lines, connecting lines, or arrows by default.
- Use small numbered markers near each target or on the floor beside each target.
- Markers must be ordered by reachable spatial depth, not randomly scattered.
- Marker 1 should usually be closest to camera; later markers move deeper into the scene.
- Each next stop must be physically reachable from the previous one without crossing walls, water, furniture, railings, people, or other obstacles.
- The scene must be designed around a walkable/glideable/flyable route before placing characters.
Only use a continuous red route line if the user explicitly asks for it or the specific workflow handles route lines reliably.
Variable Count Handling
If the user specifies a number of people/targets, build the whole asset plan, route, timeline, and video prompt around exactly that number. Do not default back to five.
Recommended target counts:
| Duration | Best count | Handling |
|---|---|---|
| 5-8s | 1-3 targets | one clear approach, few actions |
| 8-12s | 3-4 targets | short interactions, one final reveal |
| 12-15s | 4-5 targets | ideal for sequential stops |
| 15-25s | 5-8 targets | shorter beats and group transitions |
| 25s+ | 8+ targets | split into groups/zones |
If the user says “a group” without a number, choose a practical default based on duration, usually 4-5 main interaction targets. Background extras may exist only if requested and must remain secondary.
Always write: “exactly [N] main characters/targets” and “do not add or remove main targets.”
POV Identity Selection
FPV does not always mean normal human eye level. If the user specifies the POV, honor it. If not, choose a POV that makes the scene more coherent or interesting.
| Scene / intent | Good POV choices | Movement cues |
|---|---|---|
| palace, mansion, apartment, party | invited guest, butler, courier, pet cat | walking sway, room turns, hand/paw glimpses |
| modern living room, cafe, cozy home | pet cat, robot vacuum, visitor | low-angle gliding, furniture legs, paws/tail or wheel hum |
| outdoor landscape, city, battlefield | drone, bird, kite, spirit | aerial arcs, height changes, wind, no footstep bob |
| museum, exhibition, showroom | visitor, guide robot, security camera | slow walkthrough, display pauses, precise turns |
| horror / mystery | child, flashlight holder, CCTV, ghost | low height, hesitant movement, limited view |
| car/train/boat/aircraft | passenger, driver, mounted camera | constrained path, window motion, engine/road/water sounds |
POV constraints:
- Human: eye-level height, walking bob, breathing, hands/sleeves may appear.
- Pet cat/dog: low height, curious head turns, paws/tail may appear, no human hand gestures.
- Drone: smooth continuous flight, altitude changes, banking turns, propeller/wind sound, no walking bob.
- Robot vacuum/small robot: very low height, slow floor gliding, furniture legs loom large, mechanical hum, cannot climb stairs or jump.
- Bird/insect/spirit: floating arcs or gliding, unusual height, but still continuous and motivated.
- Object POV: movement must follow the object’s plausible motion unless the concept is surreal.
Always enforce physical limits. A cat cannot float over a table; a drone cannot pass through walls; a robot vacuum cannot jump onto a sofa; a human cannot squeeze through furniture.
Video Prompt Structure
Default structure:
1. Input usage: first frame, numbered route reference, character references, or red-line route-control image 2. Specs: aspect ratio, duration, one-shot 3. POV identity and physical movement 4. Route order: numbered target stops for close interaction, or geography segments for red-line path control 5. Compact timeline with one beat per target or one beat per route segment 6. Environment motion and sound design if useful 7. Consistency and avoid list
Keep the final video prompt directly copyable. Do not impose a fixed character limit unless the user asks for one.
Template: Video Prompt
使用上传图片作为首帧、编号路线参考和 [N] 个角色/目标外观参考,生成一段 [比例]、[时长] 秒、一镜到底的 [场景类型] [POV身份] FPV 视频。首帧中的编号 [1..N] 只作为镜头停靠顺序参考,不要出现在最终画面里。最终视频不要出现数字、编号、路线、箭头、文字标签或 UI。
观众/摄像机是一位/一台 [POV身份],从 [起点] 出发,严格按编号顺序移动:[起点] → 1 [目标1] → 2 [目标2] → ... → [终点/回望]。本视频包含 exactly [N] 个主要人物/目标:[列出名称],不要增加或减少主目标。角色脸、服装、位置、身份保持参考图一致。
镜头运动必须符合 [POV身份] 的物理限制:[低机位/步行/飞行/滑行/车载等具体运动];移动连续,有真实加速减速,不能瞬移、跳切、穿过障碍物或变成其他 POV。
时间轴:
0.00-[入口时间]:[建立起点和环境动态]
[时间]-[时间]:移动到 1,[目标1动作/台词/反应]
[时间]-[时间]:移动到 2,[目标2动作/台词/反应]
...
[结尾时间]:到达 [终点] 或回望全景,形成空间层次。
环境动态:[水波/窗光/咖啡蒸汽/灯光/风/反光/布料/屏幕等]。
避免:编号残留、路线/箭头、换脸、同脸、身份漂移、主目标数量变化、穿模、跳切、瞬移、错误 POV、过度抖动、场景漂移、低质、畸形、水印。Common Pitfalls
1. Asking GPT Image for a continuous red route line by default.
- Prefer numbered stop markers. Continuous lines often become disconnected, cross obstacles, or contaminate the video.
2. Letting numbered stops scatter randomly.
- First define a walkable/glideable/flyable spatial route, then place numbered stops along it.
3. Ignoring the user’s specified count.
- Use exactly the requested number in assets, route, timeline, and constraints.
4. Treating every POV like a human.
- Robot vacuums need floor gliding; cats need low curious movement; drones need flight; humans can show hands and walking sway.
5. Giving every person full dialogue in high-count scenes.
- For more than 6 targets, use groups, gestures, glances, or quick beats.
6. Forgetting to remove markers in the final video.
- Always say numbers/markers are planning references only and must not appear in final output.
7. Forgetting that red-line path control and numbered stops solve different problems.
- Numbered stops are better for characters and interiors; red-line route control is better for world-scale flythroughs, racing lines, and map-to-world journeys.
Verification Checklist
Before returning the prompt, check:
- [ ] Aspect ratio and duration are specified
- [ ] First frame / route reference / character references are assigned
- [ ] Main target count is explicit and matches the user request
- [ ] POV identity is explicit or reasonably inferred
- [ ] POV movement physics match the identity
- [ ] Route order is explicit and physically continuous
- [ ] Timeline covers the full duration and fits the target count
- [ ] Each target has an action/reaction, or group beats for high-count scenes
- [ ] Character/object consistency constraints are repeated
- [ ] Environment secondary motion is included where useful
- [ ] Negative constraints mention no route lines/arrows/numbers/markers in final video
- [ ] Negative constraints mention no added/removed main targets
- [ ] Red-line path-control mode is used only when the scene is route-shaped and large-scale, or the user explicitly requests it
- [ ] Generated prompts are in Chinese by default unless the user explicitly requested English or a target tool requires English
- [ ] Final prompt is directly copyable
References
Read references/session-patterns.md for examples and session-specific lessons from the initial development of this workflow.
Read references/mayz-seedance-world-route-case.md for the Seedance 2.0 red-line world-route pattern: aerial map control image, strict drawn route geometry, map-to-world flythrough, biome progression, speed-run/racing-line mode, and theme-park ride style journeys.
Read references/public-article-angle.md when explaining this workflow publicly in a WeChat article, X thread, tutorial intro, or demo write-up. It captures the session framing: FPV prompts are action trajectories, not just visual descriptions; numbered stops vs red-line path control; space-first design; POV physics; count-duration tradeoffs.
Read references/cafe-cat-numbered-stops-case.md when the user asks for cafe, cozy indoor, pet-cat POV, or low-angle numbered-stop interaction scenes. It captures a 15-second / 3-person cafe route pattern and the cat-specific movement constraints that prevent drone-like motion.
# OS files
.DS_Store
Thumbs.db
# Editor files
.vscode/
.idea/
*.swp
# Build / package artifacts
*.skill
*.zip
# Python cache
__pycache__/
*.pyc
Numbered Stop Marker Example
User request
帮我做一个现代客厅里 5 个人依次互动的 FPV 视频提示词,15 秒,像客人走进房间一样。
Why numbered stops
This is a close-interaction interior scene. A continuous red path can easily cross furniture, walls, or people, and the line may remain visible in the final video. Small numbered markers are safer because the prompt can define the movement between stops while the first frame only marks target order.
Asset plan
1. First-frame image of a modern living room with exactly 5 small numbered stop markers near the 5 main characters. 2. Five separate character reference images, one for each main character. 3. Optional clean first frame without numbered markers for safer image-to-video input.
Copy-ready video prompt
使用上传图片作为首帧、编号路线参考和 5 个角色外观参考,生成一段 16:9、15 秒、一镜到底的现代客厅客人视角 FPV 视频。首帧中的编号 1 到 5 只作为镜头停靠顺序参考,不要出现在最终画面里。最终视频不要出现数字、编号、路线、箭头、文字标签或 UI。
观众是一位刚进入房间的客人,从客厅入口出发,严格按编号顺序移动:入口 → 1 沙发旁的朋友 → 2 落地窗前的人 → 3 茶几旁的人 → 4 开放式吧台旁的人 → 5 阳台门口的人。全片包含 exactly 5 个主要人物,不要增加或减少主目标。角色脸、服装、位置和身份保持参考图一致。
镜头保持人类眼平高度,有轻微步行摆动、自然转头、真实加速减速和短暂停留。移动必须沿客厅地面和可通行空间完成,不能穿过沙发、茶几、墙面、玻璃、人物身体或其他障碍物。
时间轴: 0.00-2.00:从入口进入客厅,环境光从落地窗洒入,镜头建立空间方向。 2.00-4.50:移动到 1,沙发旁的朋友抬头微笑并举杯示意。 4.50-7.00:移动到 2,落地窗前的人转身看向镜头,窗帘轻微摆动。 7.00-9.50:移动到 3,茶几旁的人递出一本杂志,咖啡有轻微蒸汽。 9.50-12.00:移动到 4,吧台旁的人把杯子放下,吧台灯有温暖反光。 12.00-15.00:移动到 5,阳台门口的人拉开门,镜头停在室内外光线交界处并回望客厅层次。
避免:编号残留、路线线条、箭头、换脸、同脸、身份漂移、主目标数量变化、穿模、跳切、瞬移、错误 POV、过度抖动、场景漂移、低质、畸形、水印。
Red-Line Route-Control Example
User request
我想做一个 Seedance 2.0 的世界地图飞行,从雪原穿过峡谷、王城,最后到火山。
Why red-line path control
This is a world-scale route. The camera path is more important than close character interaction, so a red-line route-control image is useful. The red line controls geometry only and must be removed from the final video.
Route-control image prompt
生成一张 16:9 高分辨率奇幻大陆航拍路线规划图。画面从左下角雪原开始,经过中部峡谷,穿过右上方王城上空,最后抵达远处火山口。用一条清晰连续的红色路线线条从起点连接到终点,路线必须沿山谷、城门、桥梁和可飞行空域自然穿过,不要穿过厚实山体或封闭建筑。红线可以有轻微箭头指示方向,但不要出现多条路线、文字标签、复杂 UI 或过多标注。
Copy-ready video prompt
使用上传图片作为路线规划图。图中的红色路线和箭头只作为摄像机路径控制,不是最终画面内容。生成一段 16:9、12 秒、一镜到底的无人机 FPV 奇幻大陆飞行视频。镜头必须严格沿红线几何从雪原出发,穿过峡谷和王城上空,最后抵达火山口。
最终视频不要出现红线、箭头、地图标注、文字标签、UI 或俯视地图感。画面应该像真实进入这个世界,而不是停留在平面地图上。
镜头运动为隐形无人机 POV,连续向前飞行,有自然倾斜、贴近地形掠过、穿越峡谷时的前景视差、飞过王城时的高度变化和接近火山时的逐步加速。镜头不能倒退、瞬移、跳切、穿过山体或变成第三人称跟拍。
时间轴: 0.00-2.00:从雪原低空起飞,掠过冰面和风雪,沿红线方向进入山谷。 2.00-5.00:穿过峡谷,岩壁快速从两侧掠过,镜头有轻微 banking 和高度调整。 5.00-8.00:飞越王城上空,城墙、桥梁和屋顶形成明显前景视差,保持路线连续。 8.00-10.50:离开王城进入黑色火山地带,光线变暖,烟雾和火山灰增加。 10.50-12.00:抵达火山口上方,镜头减速并稳定悬停,形成终点压迫感。
避免:红线残留、箭头残留、文字标签、地图 UI、俯视平面地图感、瞬移、跳切、反向运动、穿山、穿城墙、可见无人机、错误 POV、地标变形、结构闪烁、水印。
MIT License
Copyright (c) 2026 zhouluobo
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
<div align="center">
FPV Camera Director.skill
Turn AI video prompts from visual description into action trajectory design.
   
<br>
A route-first prompting skill for cinematic FPV image-to-video workflows.
It does not only add words like cinematic, shallow depth of field, and beautiful lighting. It designs who the camera is, where it starts, which targets it visits, how it moves around obstacles, and where the shot ends.
<br>
Examples · Install · What it solves · How it works · Asset packs · 中文
</div>
---
Examples
Cat POV in a cafe
User request:
Create a 15-second FPV video prompt in a cafe where 3 people interact one by one. Cat POV.The skill recognizes this as a close character-interaction scene, not a red-line route-control scene. It prepares a full image asset pack instead of stopping at a single first-frame prompt.
Image 1: low-angle cafe first frame from cat POV, with small numbered stops 1, 2, 3
Image 2: independent reference image for the woman by the window
Image 3: independent reference image for the barista
Image 4: independent reference image for the man at the corner table
Image 5: optional clean first frame without numbers, used as the real image-to-video inputThen it produces a video prompt like this:
The viewer is a cat freely moving through a cafe. The camera stays close to the floor.
Starting from the entrance mat, it moves strictly in numbered order:
entrance mat → 1 woman by the window → 2 barista at the counter → 3 man at the corner table → stops in a patch of sunlight by the window.
The video contains exactly 3 main people. Do not add or remove main targets.
The camera movement must obey cat-body physics: low height, small steps, short pauses, curious head turns, visible table legs, chair legs, floor texture, shoes, and occasional paw or tail edges.
The cat cannot fly, jump onto the counter, pass through furniture, pass through legs, or suddenly become human eye level.The core is not making the cafe prettier. The core is making the model understand how the cat moves.
---
World-map flythrough
User request:
I want a Seedance 2.0 world-map flight, starting from a snowfield, crossing a canyon and royal city, and ending at a volcano.The skill switches to red-line route-control mode.
Image 1: 16:9 fantasy continent route-planning image, with one continuous red route from the snowfield through the canyon and city to the volcano
Image 2: optional clean world reference without the red lineThe video prompt then states:
The red route is only camera-path control, not final visual content.
The final video must not show red lines, arrows, map labels, text, UI, or a flat map-view look.
The camera must strictly follow the drawn route geometry, with natural banking, close terrain passes, foreground parallax through landmarks, and a stable horizon.Both are FPV videos, but they require different route logic.
---
What it solves
Many AI video prompts look complete and still fail.
Common failures:
- the camera teleports and breaks the space
- character count changes mid-shot
- numbers, red lines, arrows, or labels remain in the final video
- cat POV turns into drone POV
- one-shot video becomes jump cuts
- indoor movement passes through tables, chairs, walls, or people
- character faces and outfits drift between interactions
The problem is usually not a lack of style words. It is a missing action trajectory.
FPV prompting needs to describe motion, not only imagery: who is looking, how they move, whom they pass, where they pause, and what must stay consistent.
---
Use cases
- one-shot FPV walkthroughs in cafes, living rooms, galleries, courtyards, palaces, and exhibitions
- short videos where 3 to 8 characters interact in sequence
- non-human POVs such as cats, dogs, robot vacuums, drones, birds, spirits, vehicles, or object cameras
- GPT Image / GPT-Image-2 first-frame and character-reference asset packs
- Seedance, Kling, Runway, Veo, and similar image-to-video workflows
- red-line route control, world-map flythroughs, city-to-landmark movement, canyon flights, racing lines
- creators who want camera choreography instead of generic visual prompting
---
How it works
The skill breaks an FPV video into eight questions.
1. Who or what is the camera?
2. Where does it start?
3. Exactly how many main characters or targets exist?
4. In what order should the camera visit them?
5. Can every movement segment physically happen?
6. What interaction happens at each stop?
7. Which identities, outfits, and positions must remain consistent?
8. What must never appear in the final video?Then it chooses one of two route modes.
Numbered stop markers
Best for close character interactions, interiors, cafes, living rooms, social scenes, and exhibitions.
In these scenes, a red line is often risky. It can cross furniture, walls, or people, and it can leak into the final video.
The safer pattern is to place small numbered markers near each target in the first frame, then use the video prompt to define the movement between stops.
Red-line path control
Best for large-scale routes: aerial maps, city flythroughs, canyon flights, racing lines, fantasy continents, and Seedance-style path-control demos.
Here the route geometry is the main constraint. The red line can control the path, but it must disappear completely from the final video.
---
GPT Image asset packs
When the scene contains N main characters, the skill defaults to a full asset pack.
1 numbered first-frame image
N independent character reference images
1 optional clean first-frame imageFor a three-person cafe scene:
Image 1: cafe cat-POV first frame with numbered stops 1, 2, 3
Image 2: reference image for the woman by the window
Image 3: reference image for the barista
Image 4: reference image for the man at the corner table
Image 5: clean first frame without numbers or marksThe first frame controls space. The character references control identity. The clean first frame controls the actual video input.
A single image often mixes all three responsibilities and causes route marks, identity drift, or layout confusion.
---
Install
Option 1: Install into Hermes
git clone https://github.com/zhouwei713/fpv-immersive-video-prompting.git \
~/.hermes/skills/creative/fpv-immersive-video-promptingRestart Hermes or open a new session. The skill will be available as:
fpv-immersive-video-promptingOption 2: Use it in another agent runtime
If you do not use Hermes, copy SKILL.md into Claude, Codex, Cursor, OpenCode, or any agent environment that supports skills or long prompt instructions.
Option 3: Read it as a prompting method
Start with:
SKILL.md
skill/references/gpt-image-asset-packs.md
skill/references/session-patterns.md---
Usage
Try prompts like:
Create a 15-second cat-POV FPV video in a cafe where 3 people interact one by one.Make a palace courtyard FPV shot that passes 5 characters in order and ends by a pond.I have a world map. Help me write a Seedance 2.0 prompt where the camera follows a red route from a snowfield to a volcano.Give me a GPT Image asset-pack prompt with a first frame, character references, and a clean first frame.---
Repository structure
.
├── README.md
├── README_EN.md
├── SKILL.md
├── skill/
│ ├── SKILL.md
│ └── references/
│ ├── gpt-image-asset-packs.md
│ ├── liyue-ai-redline-fpv-case.md
│ ├── mayz-seedance-world-route-case.md
│ ├── public-article-angle.md
│ └── session-patterns.md
├── examples/
│ ├── numbered-stop-example.md
│ └── redline-route-example.md
└── LICENSE---
Design philosophy
FPV prompting is not about making one still image more beautiful.
It is closer to level design: where the entrance is, how the route moves, where the characters stand, whether the camera body can physically pass through the space, and whether the ending gives the viewer a coherent sense of the scene.
If motion is not designed, the model designs it for you. That usually means teleportation, wall-crossing, face drift, and jump cuts.
This skill makes the route, identity, and physics checks explicit and reusable.
---
License
MIT License.
<div align="center">
FPV 运镜导演.skill
把 AI 视频提示词,从“画面描述”升级成“行动轨迹设计”。
   
<br>
一个专门为 FPV / image-to-video 设计的 AI 视频提示词 Skill。
它不只帮你写“电影感、浅景深、光影高级”。 它会先设计镜头是谁,从哪里出发,按什么顺序经过哪些人,怎么绕过障碍物,最后停在哪里。
<br>
看效果 · 安装 · 它解决什么 · 工作原理 · 资产包 · English
</div>
---
效果示例
咖啡厅猫咪视角
用户输入:
帮我做一个咖啡厅里 3 个人依次互动的 FPV 视频提示词,15 秒,猫咪视角Skill 会先判断这不是红线路径场景,而是近距离人物互动。它会生成一套完整资产包,而不是只给一张首帧图。
图片 1:咖啡厅猫咪低机位首帧,带 1、2、3 小编号停靠点
图片 2:靠窗女生独立人物参考图
图片 3:吧台咖啡师独立人物参考图
图片 4:角落圆桌男生独立人物参考图
图片 5:干净版首帧,去掉编号,用作真正视频首帧然后输出视频 prompt:
观众是一只在咖啡厅里自由走动的猫咪,镜头保持接近地面的低机位,
从咖啡厅入口旁的地垫开始,严格按编号顺序移动:
入口地垫 → 1 靠窗座位的女生 → 2 吧台前的咖啡师 → 3 角落圆桌旁的男生 → 窗边阳光下停住。
全片包含 exactly 3 个主要人物,不要增加或减少主目标。
镜头运动必须符合猫咪的身体限制,能看到桌腿、椅腿、地面纹理、人的鞋子,
可以短暂停顿、好奇转头、绕开障碍物,但不能飞,不能跳上吧台,不能穿过桌椅和人腿。这类提示词的重点不是“咖啡厅很漂亮”,而是让视频模型知道猫到底怎么走。
---
世界地图飞行
用户输入:
我想做一个 Seedance 2.0 的世界地图飞行,从雪原穿过峡谷、王城,最后到火山。Skill 会切换到红线路径控制模式。
图片 1:16:9 奇幻大陆航拍路线规划图,一条连续红线从雪原出发,经过峡谷和王城,抵达火山口
图片 2:可选干净世界参考图,去掉红线视频 prompt 会明确说明:
红色路线只作为摄像机路径控制,不是最终画面内容。
最终视频不要出现红线、箭头、地图标注、文字标签、UI 或俯视地图感。
镜头必须严格沿红线几何飞行,有自然 banking、贴近地形掠过、穿越地标时的前景视差和稳定地平线。同样是 FPV,一个是猫在咖啡厅里走,一个是无人机穿越大陆。两种场景不能用同一套提示词。
---
它解决什么
很多 AI 视频 prompt 看起来很完整,实际一生成就翻车。
常见问题是:
- 镜头突然瞬移,空间断了
- 人物数量一会儿多一会儿少
- 首帧里的编号、红线、箭头残留在成片里
- 说是猫咪视角,结果变成无人机视角
- 说是一镜到底,结果中间跳切
- 室内路线穿过桌子、椅子、墙和人腿
- 多人物互动里,角色脸和衣服互相漂移
原因通常不是缺少风格词,而是缺少行动轨迹。
FPV 视频要写清楚的不是一张图,而是一段运动:谁在看,怎么走,经过谁,在哪里停,什么东西不能变。
---
适合什么场景
- 咖啡厅、客厅、展厅、庭院、宫殿等室内一镜到底
- 3 到 8 个角色依次互动的短视频
- 猫、狗、机器人吸尘器、无人机、鸟、幽灵、车辆等非人类 POV
- GPT Image / GPT-Image-2 首帧和人物参考图资产包
- Seedance、Kling、Runway、Veo 等 image-to-video 工作流
- 红线路径控制、世界地图飞行、城市到地标、峡谷穿越、赛车线路
- 想把提示词从“视觉描述”变成“镜头调度”的创作者
---
工作原理
这个 Skill 会把一个 FPV 视频拆成 8 个问题。
1. 摄像机是谁
2. 从哪里开始
3. exactly 有多少个主要人物或目标
4. 按什么顺序经过它们
5. 每一段路线是否物理可达
6. 每个停靠点发生什么互动
7. 哪些身份、服装、位置必须保持一致
8. 哪些东西绝对不能出现在最终画面里然后它会选择两种路线模式之一。
编号停靠点
适合近距离人物互动、室内空间、咖啡厅、客厅、派对、展厅。
这类场景里不要默认画红线。红线很容易穿过桌椅、墙面和人腿,也容易残留在视频里。
更稳定的做法是,在首帧里放小编号 1、2、3,把角色顺序标清楚。真正的移动路线交给视频 prompt 约束。
红线路径控制
适合大世界路线、航拍地图、城市飞行、峡谷穿越、赛车线路、Seedance 2.0 path control。
这类场景的核心是路线几何。红线可以作为路径控制,但最终视频里必须完全消失。
---
GPT Image 资产包
如果场景里有 N 个主要人物,Skill 默认生成完整资产包。
1 张带编号首帧图
N 张人物独立参考图
1 张可选干净首帧图以 3 人咖啡厅为例:
图片 1:咖啡厅猫咪视角首帧,带编号 1、2、3
图片 2:靠窗女生参考图
图片 3:咖啡师参考图
图片 4:角落男生参考图
图片 5:干净版首帧,去掉编号和所有标记这样做的目的很简单:首帧管空间,人物参考管身份,干净首帧管最终输入。
如果只给一张图,视频模型很容易把路线、角色和画面标记混在一起。
---
安装
方式一:安装到 Hermes
git clone https://github.com/zhouwei713/fpv-immersive-video-prompting.git \
~/.hermes/skills/creative/fpv-immersive-video-prompting重启 Hermes,或开启新会话后即可使用:
fpv-immersive-video-prompting方式二:作为通用 Skill 使用
如果你不使用 Hermes,也可以直接把 SKILL.md 放进 Claude、Codex、Cursor、OpenCode 或其他支持 Skill / long prompt 的 Agent 环境里。
方式三:只当提示词方法论参考
直接阅读:
SKILL.md
skill/references/gpt-image-asset-packs.md
skill/references/session-patterns.md---
使用方式
你可以这样说:
帮我做一个咖啡厅里 3 个人依次互动的 FPV 视频提示词,15 秒,猫咪视角做一个宫殿庭院 FPV,一镜到底经过 5 个角色,最后到池塘边我有一张世界地图,想让镜头沿红线从雪原飞到火山,帮我写 Seedance 2.0 prompt给我一套 GPT Image 生图 prompt,要包含首帧图、人物参考图和干净版首帧---
仓库结构
.
├── README.md
├── README_EN.md
├── SKILL.md
├── skill/
│ ├── SKILL.md
│ └── references/
│ ├── gpt-image-asset-packs.md
│ ├── liyue-ai-redline-fpv-case.md
│ ├── mayz-seedance-world-route-case.md
│ ├── public-article-angle.md
│ └── session-patterns.md
├── examples/
│ ├── numbered-stop-example.md
│ └── redline-route-example.md
└── LICENSE---
设计理念
FPV 提示词不是把一张图说得更美。
它更像在设计一个小关卡:入口在哪里,路线怎么走,角色站在哪里,镜头能不能绕过去,最后有没有一个让观众理解空间的停顿。
只要运动没设计清楚,模型就会替你设计。它设计出来的结果,通常就是瞬移、穿墙、换脸和跳切。
这个 Skill 的价值,是把这套判断流程固定下来,让每次生成 FPV 视频前都先过一遍空间、路线、角色和物理限制。
---
English
For English documentation, see README_EN.md.
---
License
MIT License.
GPT Image Asset Packs for FPV Video
Use this reference when a user wants GPT Image 2 / ChatGPT image generation to prepare assets for an FPV image-to-video workflow.
Multi-image asset pack beats single contact sheet
For ChatGPT / GPT Image workflows, deliver one batch-generation prompt that asks the model to generate multiple separate images in a single response. Do not deliver many independent prompts unless the user explicitly asks for per-image prompts. The prompt should still request separate images, not one crowded contact sheet.
If ChatGPT can generate multiple images in one response, request separate images. Do not stop at the first-frame prompt when the scene contains named people/targets; identity consistency needs independent reference images.
For close-interaction / character scenes, request:
1. Numbered first frame
- 16:9 scene from the intended starting viewpoint
- uses small numbered stop markers near each main target
- no red line unless the user explicitly asks for a drawn route
2. Character references
- one image per main character/target
- clean portrait or full-body card
- no labels, no route marks, no text if the video model is sensitive to text
3. Optional clean first frame
- same composition as the numbered first frame
- removes numbers, arrows, route marks, and labels
- best used as the actual image-to-video first frame
For world-route / path-control scenes, request:
1. Red-line first frame or route-control map
- 16:9 route-planning image from the intended starting viewpoint or aerial map
- includes one continuous red route line and ordered route geometry
2. Optional clean world/scene reference without the red line
3. Optional landmark or character references if specific destinations need stable appearance
Video prompt rule: numbered markers and red-line marks are route-planning references only and must not appear in the final video.
Continuous-route wording
Use topological constraints. GPT Image often places 1,2,3,4,5 as decorative labels unless the prompt explicitly defines a walkable route.
Reliable language:
This image is not a random group scene. First design one continuous walkable space from foreground to background. Draw a single unbroken red route line on the floor/path. The line starts at the camera entrance, follows only physically walkable ground, and passes through stops 1, 2, 3, 4, 5 in order. Each next stop must be physically reachable from the previous stop without crossing walls, water, railings, furniture, or people. Stop 1 is closest to camera, stop 2 is slightly farther, stop 3 is midground, stop 4 is deeper in the scene, and stop 5 is the destination/final reveal. Do not scatter the stop numbers around the image. Do not draw multiple route lines or disconnected segments.Adapt this for non-human POVs:
- drone: “continuous flyable air corridor”, avoid walls/ceilings/people
- pet cat: “low floor-level path”, route goes under/around furniture but never floats
- robot vacuum: “floor-only path”, cannot climb stairs or cross thick rugs if unrealistic
- bird/spirit: “continuous aerial arc”, still no teleporting or passing through solid objects
Count-aware asset request pattern
If the user specifies N people/targets in a close-interaction or character scene, request exactly:
- 1 numbered first-frame image containing exactly N main stops
- N character reference images
- optionally 1 clean first-frame image
If the user specifies N people/targets in a world-route/path-control scene, request exactly:
- 1 red-line route-control image or route-planning map
- N character/landmark references only when those targets must stay visually consistent
- optionally 1 clean route/world reference without red line
Include this line:
Generate exactly [N] main characters/targets and exactly [N] stop points. Do not add extra main characters. Background extras are allowed only if I explicitly ask for a crowd, and they must remain visually secondary.Example asset-pack skeleton
请一次性生成 [N+1] 张图片,用于后续制作一段 [场景] 第一人称 FPV 沉浸式 AI 视频。所有图片保持同一套美术风格、同一空间设定、同一光影质感、同一人物设计语言。
图片 1:[场景] FPV 首帧图,带连续红色运镜路线。
生成一张 16:9 横版图片,摄像机身份是 [POV],从 [起点] 看向 [空间]。这张图的核心是一条清晰、连续、物理可行的路线。请先设计一个从前景到远景逐步深入的空间:[列出前景、近景、中景、远景]。
在 [地面/空中/道路/水面] 上画一条单一、连续、不断开的红色运镜路线,从 [起点] 出发,依次经过 1 到 [N] 个停靠点,最后到达 [终点]。每个点必须在同一条路线上,按行进顺序排列,不要随机散布。路线不能穿过 [当前场景的障碍物]。
[N] 个主要目标沿这条连续路线依次分布:
1 [名称]:[位置],[外观],[动作/状态]
2 [名称]:[位置],[外观],[动作/状态]
...
图片 2 到 图片 [N+1]:分别生成 [N] 个主要角色/目标的独立参考图。每张参考图背景简洁,不要文字、标签或红线。角色脸部、服装、颜色和气质要明显不同。
如果系统允许额外生成 1 张,请生成干净版首帧图,内容与图片 1 一致,但去掉红色路线、箭头、编号和所有文字标记。Session lessons captured
- A single preproduction board is less ideal than multiple separate images when ChatGPT image generation can output multiple images.
- Red-line first frames are acceptable if the later video prompt explicitly forbids showing the red line in final output.
- Route continuity must be specified as topology, not just as “1→2→3→4→5”.
- User-specified target count must drive number of reference images, stops, timeline beats, and negative constraints.
- POV identity changes the entire route physics; write for the selected camera body, not generic “FPV”.
Li Yue red-line FPV / Seedance 2.0 case note
Source pattern: an X thread by @liyue_ai about a 15-second immersive ancient-style FPV game-demo video made with GPT Image 2 + red-line camera route planning + Seedance 2.0.
What matters
The value is not the one long prompt itself. The reusable pattern is:
1. Generate a static scene / first frame with clear spatial layout. 2. Generate or provide separate character references so the video model can maintain faces and clothing when the camera focuses on each person. 3. Draw a red camera path over the scene as a planning artifact. 4. Convert the path into a timestamped route. 5. Give every stop on the route a small interaction beat: position, action, dialogue, info card, consistency constraints. 6. Add first-person body physics: walking sway, breathing rhythm, small head-bob, acceleration/deceleration, foreground/background parallax. 7. Add subtle environmental motion: water, fabric, curtains, hair, light, reflection. 8. State that red path lines/arrows are only for planning and must not appear in the output. 9. Add an explicit avoid list for identity drift, extra characters, morphing, route-line leakage, camera teleporting, clipping, excessive shake, and UI covering faces.
Example interaction structure
For each character/object:
[time range]: The player approaches [position + identity]. [Character/object] does [small action]. [Dialogue/sound/UI card]. Keep [face/clothing/position/count] consistent.Ancient-garden sample beat map
- 0.00–0.70s: establish first-person entry, water ripples, curtains, ambience.
- 0.70–2.30s: approach 清扬, purple seated character, singing specialty, greeting line and info card.
- 2.30–4.20s: approach 静姝, green seated character with fan, chess specialty.
- 4.20–6.40s: approach 令仪, central pink standing character, qin specialty.
- 6.40–8.70s: approach 杜若, blue character by railing, calligraphy/writing specialty.
- 8.70–12.20s: move through pavilion, emphasize parallax; background characters only breathe/blink.
- 12.20–14.50s: approach 采薇, pink looking-back character, painting specialty.
- 14.50–15.00s: settle into final layered group composition.
Why comments matter
For posts like this, comments often contain the real workflow details: tool names, platform names, route-planning trick origin, whether arrows are only planning guides, model/version, and constraints discovered from failures. When researching similar examples, inspect the author’s replies and follow-up posts, not just the main media post.
Reusable output rule
When turning a viral post into a skill, avoid making a one-case skill like “Li Yue ancient garden prompt”. Generalize it into the class: first-person route-planned immersive video prompting. Keep the specific post as a reference file.
Mayz / Seedance 2.0 world-route FPV case note
Source pattern: an X thread by @Mayz1169 about Seedance 2.0 path control, using a Lord-of-the-Rings-style aerial terrain map with one red route line to generate a 15-second cinematic FPV flight.
What the post demonstrates
The reusable trick is different from close-range character-stop FPV:
1. Create a top-down or high aerial world map / terrain route image. 2. Draw one clear red line showing the entire camera path and direction. 3. Feed the map into Seedance 2.0 as a route-planning control image. 4. In the video prompt, tell the model the red line is only the intended flight route and must be completely removed from the final video. 5. Convert the route into a timed world progression: peaceful zone → transition zone → city/fortress → hostile zone → destination reveal. 6. Use camera-language constraints: invisible first-person drone, no map view, no cuts, no teleporting, strict route geometry, natural banking, altitude changes, progressive acceleration, foreground parallax. 7. Use environment progression as the story engine, not character dialogue.
When to use this mode
Use red-line path control when the scene is large-scale and route-shaped:
- fantasy continent journey
- game trailer / open-world flythrough
- city-to-landmark FPV
- map-to-world transformation
- racing path / canyon flight / aerial drone route
- theme-park ride style journey
- product/world showcase with clear geography
Do not use this as the default for indoor scenes, close character interactions, or crowded social scenes. For those, numbered stop markers are usually safer because red route lines can cross people/furniture or leak into the output.
Prompt ingredients that matter
Route image prompt
Ask for:
- high-resolution 16:9 aerial terrain map or world map
- clear start and destination regions
- one continuous red route line, optionally with a subtle arrow
- route drawn through physically plausible corridors: roads, valleys, rivers, city gates, bridges, canyons, coastlines, rooftops, airspace
- distinct visual zones along the path so the video has progression
Video prompt pattern
Use the uploaded image as a route-planning terrain map. The red line indicates the intended camera flight path and direction only. The red line, arrow, and all annotations must be completely removed from the final video.
Create a [duration]-second single continuous [aspect ratio] cinematic FPV flight through the exact world shown in the image. No cuts, no transitions, no teleporting, no map view. The camera is an invisible first-person drone camera and must strictly follow the drawn route geometry from [start] to [destination].
Timeline:
0-[t1]s: [zone 1, low close passes, parallax]
[t1]-[t2]s: [zone 2, route-following transition]
[t2]-[t3]s: [major landmark / city / obstacle]
[t3]-[t4]s: [dramatic terrain change]
[t4]-[end]s: [destination reveal / climb / final scale shot]
Camera motion: fast cinematic FPV, continuous forward flight, natural banking, close passes, dynamic altitude changes, strong foreground parallax, progressive acceleration, smooth horizon control.
Avoid: visible red line, visible arrows, annotations, text, subtitles, logos, watermarks, map appearance, jump cuts, teleportation, reverse movement, visible drone, guide characters, modern/sci-fi objects unless requested, blurry terrain, deformed buildings, flickering structures, flat environments.New gameplay patterns enabled by this case
1. World-transition ride
- The route itself is the story: safe village → capital city → ruins → volcano / boss arena.
- Best for trailers, lore intros, game-world showcases.
2. Mission-route briefing becomes gameplay footage
- First image looks like a tactical map, final video becomes the actual FPV route.
- Good for spy infiltration, fantasy quest, heist, battlefield flythrough.
3. Biome progression
- Each route segment changes weather, lighting, architecture, and danger level.
- Useful for showing model control over long visual transitions.
4. Landmark chain
- The route links 4-6 large landmarks rather than 4-6 people.
- Works better with drone/bird/spirit POV than human walking POV.
5. Map-to-world transformation
- The control image can be a stylized map, but the prompt must say final output is not a map view; it becomes real cinematic terrain.
6. Speed-run / race-line mode
- Use the red path as a racing line through canyon, city rooftops, tunnel, bridge, forest, or sci-fi trench.
- Add speed cues, banking, near misses, motion blur, and checkpoint reveals.
7. Theme-park dark ride
- A continuous guided ride through a designed world: entrance → story zones → threat reveal → finale.
- Better if the route has curves, gates, tunnels, and reveal moments.
Practical boundary
This case weakens the old rule “avoid red lines by default” only for world-scale route control. Keep both modes:
- Numbered stops: close interaction, characters, indoor scenes, social scenes, exact target count.
- Red-line path control: large terrain, aerial route, long-distance journey, continuous geography, environment progression.
If the user asks for Seedance 2.0 path control, drawn route, game-map flight, or world traversal, choose red-line path mode by default.
Public article angle: action trajectory as the core of FPV video prompts
This reference captures the public-facing explanation of the FPV immersive video prompting workflow after writing a WeChat article from the skill.
Article-level framing
The strongest general-audience angle is:
AI video prompts are not only about visual quality. For FPV / image-to-video scenes, the most fragile part is the action trajectory: who/what the camera is, where it starts, which stops it visits, how each segment is physically reachable, and where it ends.
Use this framing when the user asks for an explanatory article, X thread, tutorial intro, or demo write-up about this skill.
Key narrative lessons
1. Treat FPV video as a small playable scene or game level.
- Start point, route, stop order, target count, POV physics, and final reveal matter as much as style words.
2. For indoor / close character interaction scenes, numbered stop markers are usually easier to explain than a continuous red route line.
- A bad red line can cross furniture, water, railings, people, or walls and then mislead the video model.
- Numbered markers can be paired with strong prompt language: move from 1 to 2 to 3 in order, follow visible floor/corridor paths, no teleporting, no cuts, no obstacle crossing.
3. For world-scale routes, red-line path control remains useful.
- Fantasy continent flythroughs, city-to-landmark routes, racing lines, canyon flights, world maps, and Seedance 2.0 path-control demos are better candidates for red-line route images.
4. Design the space before placing characters.
- Example: for a modern living room, first define an open route from entrance → sofa → window → coffee table → bar → balcony, then place people along that route.
- If characters are placed first as a beautiful group portrait, route continuity often breaks.
5. Count, duration, and interaction density must be linked.
- 15 seconds with 5 people can support short individual beats.
- 8 or 12 people should become zones/groups or quick gestures, not full dialogue for everyone.
6. POV identity can make or break the route.
- Human: eye height, walking bob, hands/sleeves.
- Cat: low height, furniture legs, paws/tail, no flying.
- Robot vacuum: floor-only gliding, cannot climb stairs or jump.
- Drone/bird/spirit: aerial arcs, banking, altitude changes, no footstep bob.
Reusable quote-style lines
- FPV video prompts should describe action, not only images.
- If the model does not know how the camera moves, it will invent movement for you.
- Skill value is not storing one universal prompt; it is storing the judgment process: numbered stops or red line, how many targets, which POV, what timeline, what constraints.
When writing examples
For public articles, avoid dumping full giant prompts unless the user explicitly wants templates. It is usually better to explain the decision logic with concrete scene snippets:
- modern living room: entrance → sofa → window → tea table → bar → balcony
- palace courtyard: gate → corridor column → flower tree → central steps → side corridor → pond
- fantasy map: snowfield → wall → valley → capital → strait → volcano
Then mention that the actual Skill can generate the full GPT Image asset-pack prompt and video prompt from these decisions.
Session Patterns: FPV Immersive Video Prompting
This reference captures practical lessons from developing the FPV prompt workflow.
Numbered markers beat red route lines
Initial approach used red camera-path lines on the first frame. In practice, GPT Image often struggles to draw one continuous physically coherent route. Lines may disconnect, cross obstacles, or become decorative. For most workflows, ask GPT Image for numbered stop markers only: 1, 2, 3, etc.
Video prompt pattern:
首帧中的 1、2、3 只代表镜头停靠顺序,最终画面不要出现数字、编号、路线、箭头、文字或 UI。镜头按编号连续移动:入口 → 1 → 2 → 3 → 终点。每段移动必须沿可见可行走空间自然前进,不能瞬移、跳切或穿过障碍物。Spatial ordering for markers
Do not let markers scatter across a beautiful scene. First design the environment as a route, then place targets along it:
- 1: closest to camera / foreground
- 2: near-midground
- 3: midground
- 4: mid-far / side corridor
- 5: far destination
Each next marker must be reachable without crossing walls, furniture, water, people, railings, or impossible height changes.
Variable count rule
If the user says “3 people,” use exactly 3 main characters in image prompts, route, timeline, and negative constraints. Do not return to a 5-character default. If the user says “a group” but gives no number, pick a practical count for duration, usually 4-5 main targets.
POV examples
Coffee shop, 3 people, robot vacuum POV
Asset prompt should request:
- 16:9 first frame from ~10 cm height
- numbered stops only: 1 near window customer, 2 central table customer, 3 barista at counter
- separate reference images for each of the 3 characters
- optional clean first frame without numbers
Video constraints:
- exact 3 main characters
- robot vacuum height ~10 cm
- slow floor gliding, slight mechanical vibration/turning hesitation
- cannot fly, jump, climb, or pass through table legs, chair legs, people, counter, walls, cables
- avoid accidentally turning into human eye-level or drone POV
Useful beat structure for 12s:
- 0-1s: start from entrance floor
- 1-4s: marker 1 interaction
- 4-7s: marker 2 interaction
- 7-10s: marker 3 interaction
- 10-12s: pause/turn/reveal scene layers
Do not impose a default 2000-character limit
The user briefly considered a 2000-character video prompt limit, then explicitly reverted it. Keep prompts directly usable and concise enough for the target tool, but do not bake in a fixed default character limit unless the user asks for one in the moment.