
Video
- 131 installs
- 121 repo stars
- Updated August 4, 2026
- smixs/visual-skills
Helps with ai & agent building tasks.
About
video is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted development.
- video
- AI & Agent Building
- AI-coding skill
Video by the numbers
- 131 all-time installs (skills.sh)
- +21 installs in the week ending Jul 27, 2026 (Skillselion tracking)
- Ranked #3,619 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/smixs/visual-skills --skill videoAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 131 |
|---|---|
| repo stars | ★ 121 |
| Last updated | August 4, 2026 |
| Repository | smixs/visual-skills ↗ |
What it does
Helps with ai & agent building tasks.
Files
AI Director, Screenwriter & Editor
Hybrid role. You direct (see frame, emotion, motivated camera), write (build beat, action, consequence, final image), and edit (cut rhythm, protect continuity, drive montage). Prompt engineering is fourth — it serves the first three.
A beautiful frame without dramaturgy is wallpaper. A dramaturgically clean prompt without details is mush. The whole craft of this skill lives in the reference files. The body of this SKILL.md is intentionally thin so you cannot fake a result by reading it alone.
---
Mandatory reading order — DO NOT WRITE A PROMPT WITHOUT THIS
Past attempts to write prompts directly from this skill body produced lazy, mush-prone results. The fix is structural: the process lives only in the reference files, and you load them in this order before producing output. Skipping a step silently degrades the result — the model cannot tell that a shot is wallpaper, only the writer can, and only by applying the rules from these files.
For every video prompt request, load the files in this order:
Step 1 — always read first → dramaturgy.md
Scene formula. Details Law (the second core law, most violated). Murch Rule of Six. Three-jobs rule. Five anchors. Blocking, staging, environment as pressure. Three-layer storyboard. 14-field shot card. Rhythm ladder. Dramaturgy check.
You cannot decide whether a prompt is ready without running the dramaturgy check from this file.
Step 2 — always read second → universal-rules.md
U1–U12 universal rules that apply to every video model: prompt skeleton, weight-at-start, show-don't-tell, lens language, character anchor, contradictions, duration discipline, final image rule, three-detail check.
Step 3 — pick the model and read one model file
Use this short selector. The full reasoning is in the chosen file.
| Cue from the user / task | Read |
|---|---|
Seedance, ByteDance, Doubao, multi-shot in one clip, --resolution, --duration, --camerafixed, "Cut to", @img1, fast multi-shot drama | seedance.md |
| Kling, Kuaishou, Element Binding, Motion Brush, Motion Control, dedicated negative prompt field, Kling 3.0 multi-shot with `[Character A: ...]` labels, native dialogue + lip-sync, 15s | kling.md |
| Veo, Google video, dialogue / lip-sync, JSON prompts, synchronized SFX, commercial polish with voiceover | veo.md |
Default if nothing in the request hints at a model:
- Multi-shot narrative or fast montage drama → Seedance, or Kling 3.0 if dialogue is involved.
- Dialogue / commercial polish / synchronized SFX → Veo, or Kling 3.0 for multi-character dialogue scenes up to 15s.
- Character consistency across many social clips → Kling 2.6 Pro (cheaper) or Kling 3.0 (with in-prompt
[Character A: ...]labels). - 10-15s continuous narrative with audio → Kling 3.0.
For a more detailed comparison (max clip length, audio support, character lock methods, motion brush, etc.), read the model file you picked. Do not load all three.
Step 4 — task-shaped reading (load only those that match)
- Storyboard / shot list / director treatment / "разбей на склейки" → role-modes.md. Determines whether you operate as Director, Screenwriter, or Editor for this turn.
- Commercial, music video, drama, action, fashion, UGC, product film, escalation / anxiety / discovery / catastrophe / product-drama montage → patterns-and-genres.md.
- Multi-clip continuity, fixing a broken prompt, known failure modes (one-take, face drift, melted hands, dialogue too fast) → fixes-and-skeletons.md.
- Need precise framing / lens / movement / light / sound terms → camera-lighting-vocabulary.md.
If none match — proceed with steps 1-3 only.
Step 5 — apply the dramaturgy check and the three-detail check
Before returning anything, run both checks:
- Dramaturgy check (
dramaturgy.md§15): scene formula complete, three-detail check on every shot, three-jobs rule on every shot, motivated camera, readable geometry, five anchors named. - Three-detail audit (
universal-rules.md§13): each shot owns environmental pressure + physical micro-action + sound or visual motif.
If any shot fails, fix before sending. This is the step the user has had to enforce repeatedly. Do not skip it.
---
Output
Choose the format the request actually asks for. Default to A if unclear.
- A. Single prompt. One ready-to-copy prompt for one generation. Lead with model name + parameters in a short header.
- B. Multi-clip prompts. Sequence of self-contained prompts, each repeating the full identity / style / continuity block (see
universal-rules.mdU7). - C. Storyboard. Table — Time, Shot, Function, Action, Camera, Light, Sound, Emotion. Every row is a 14-field shot card from
dramaturgy.md§11, compressed. - D. Prompt audit. Given a user prompt, return: What works, What breaks generation, Missing direction, Continuity risks, Model-specific mismatches, Stronger version (rewritten prompt).
- E. Director treatment. Core idea, Emotional arc, Visual motif, Rhythm, Camera language, Lighting, Sound, Ending image. (Treatment ≠ prompt.)
- F. JSON (Veo only). Structured scene-by-scene continuity. See
veo.md.
Default output language follows the user. The final AI prompt itself goes in English unless the user asks otherwise — Seedance, Kling, and Veo all perform better in English.
---
Final response style
Prefer: ready-to-copy prompts, clear section labels, production language, motivated camera and light direction, strict continuity blocks, model-specific syntax, direct fixes.
Avoid: long theory unless asked, academic lectures, vague inspiration, decorative jargon, "cinematic masterpiece" filler, prompts without camera and light, prompts without continuity, stacking more than two director references, abstract emotions without physical translation.
When in doubt about a model-specific detail — re-read the model file before writing the final prompt. It costs nothing and prevents bad output.
Camera, lighting, color, sound vocabulary
Use precise production language. "Cinematic" is not a direction. "35mm, slow push-in, warm window key from frame-left" is.
Contents
1. Framing 2. Camera movement 3. Lens language 4. Light sources 5. Light direction 6. Light quality 7. Color discipline 8. Sound categories 9. Blocking language
---
1. Framing
- extreme wide shot
- wide shot
- medium wide
- medium shot
- medium close-up
- close-up
- extreme close-up
- macro insert
- over-the-shoulder
- POV
- profile
- silhouette
- low angle
- high angle
- top-down
- dutch angle
- locked-off frame
2. Camera movement
- static camera
- slow push-in
- fast push-in
- pull-back
- tracking shot
- lateral tracking
- handheld micro-shake
- whip pan
- snap zoom
- rack focus
- tilt up
- tilt down
- orbit
- gimbal glide
- dolly-in
- dolly-out
- crane shot
- aerial shot
- Hitchcock zoom (dolly-zoom)
Universal rule. Pick one dominant move per 5-second shot. Layer one subtle secondary move at most.
3. Lens language
- 24mm. Wide, immersive, exaggerated proximity.
- 35mm. Natural documentary.
- 50mm. Intimate, human perspective.
- 85mm. Portrait, compressed background.
- 100mm macro. Texture, detail.
- anamorphic 40mm. Cinematic widescreen.
- shallow depth of field.
- deep focus.
- compressed telephoto background.
- distorted wide-angle proximity.
4. Light sources
- cold refrigerator light
- harsh overhead kitchen LED
- moonlight through window
- neon sign spill
- phone screen glow
- car headlights
- streetlamp backlight
- fluorescent office light
- practical lamp
- candlelight
- monitor glow
- emergency red light
- sodium vapor street light
- police strobe
- theater house lights
- stage spotlight
- campfire flicker
- sunset through blinds
- under-counter kitchen LED
5. Light direction
- top light
- side light
- backlight
- underlight
- frontal soft light
- rim light
- bounce fill
- window key light
- motivated practical (the light source is visible in frame)
6. Light quality
- hard shadow
- soft diffused
- low-key lighting
- high contrast
- silhouette
- specular highlights
- volumetric haze
- steam catching backlight
- wet reflections
- desaturated palette
- cold blue-gray grade
- teal shadows
- clean commercial lighting
- gritty naturalistic
- blown-out overexposure
- crushed blacks
7. Color discipline
Define palette with concrete colors. Never write "cinematic colors."
Bad. "cinematic colors." Good. "Cold blue-gray shadows, desaturated skin tones, greenish fridge spill, black negative space, no warm yellow tones."
If the user bans a color, obey it strictly. Repeat the ban in every clip of a multi-clip sequence.
Common palettes.
- Fincher cold. Cold blue-gray shadows, desaturated skin, black negative space.
- Deakins natural. Warm amber interiors, cool blue exteriors, clean contrast.
- Wong Kar-wai. Saturated warm reds, deep greens, hazy practicals.
- Glazer neon. Black, one dominant neon hue (magenta or teal), hard edges.
- Commercial creamy. Warm creams, soft pastels, clean whites, no harsh blacks.
- Safdie chaotic. Mixed sources, overlapping color temperatures, urban neon spill.
8. Sound categories
Even if the model does not generate audio, sound description helps structure rhythm.
Ambient.
- room tone
- distant city ambience
- rain on tin roof
- wind through leaves
- fridge hum
- fluorescent buzz
- traffic wash
Body and action.
- breath
- footsteps
- fabric rustle
- fork clink
- door hinge
- key in lock
- zipper
- stomach growl
Dramatic events.
- sudden silence
- low bass hit
- wet thud
- glass break
- distant thunder
- car door slam
- final sound cue before cut to black
For Veo, wrap sound in syntax. Audio:, SFX:, or (parenthetical description). For Seedance 1.5+, include sound in the prompt body. For Kling, sound descriptions help rhythm planning but audio is not generated.
9. Blocking language
Describe physical movement with six inputs.
- Who moves.
- Where they start.
- Where they end.
- What object they touch.
- What they look at.
- What their body reveals emotionally.
Example.
The man stands in the kitchen doorway, shoulders collapsed. He slowly approaches the refrigerator, opens it with hesitation, leans into the cold light, then freezes when he sees the empty shelves.This beats "a sad man goes to the fridge" because it is playable by the model frame by frame.
Dramaturgy, detail, montage
This is the mandatory layer. A beautiful frame without dramaturgy is wallpaper. Every prompt built by this skill must pass the dramaturgy check before it is sent to the user.
Contents
1. Core law. Scene formula 2. Second law. Details intensify emotion 3. The three-jobs rule (what every shot must do) 4. Walter Murch Rule of Six 5. Blocking as choreography of desire 6. Staging controls subtext 7. Camera must have a reason 8. Spatial clarity beats montage hysteria 9. Environment plays 10. Three-layer storyboard method 11. Shot card template (14 fields) 12. Rhythm ladder 13. One-anchor principle 14. Worked example 15. Dramaturgy check before sending any prompt
---
1. Core law. Scene formula
A scene exists only when all five elements are present.
Scene = hero's desire + obstacle + space geometry + controlled gaze + editing rhythmIf any element is missing, the scene collapses into decoration.
- Desire. What does the character want right now in this specific second.
- Obstacle. What blocks them. Object, person, fear, distance, rule.
- Space geometry. Who stands where. Who has the power position. Which direction is threat, which is escape.
- Controlled gaze. Where the viewer's eye is forced to look, one focal point per frame.
- Editing rhythm. How long each shot lives, where the pause lands, where the cut bites.
Before writing a prompt, name each of these in one sentence. If you cannot, the scene is not ready.
2. Second law. Details intensify emotion. Laziness kills the prompt.
The scene formula tells you what a scene is. The Details Law tells you how every shot must be written. Skip it and even a perfect dramaturgical structure produces mush on screen.
Every shot owns three concrete physical details: one environmental pressure, one physical micro-action, one sound or visual motif anchor.
The three-detail rule
For each shot in the storyboard or prompt, before it is sent, the writer commits to:
1. Environmental pressure. A physical fact about the space that carries the emotion. Cold refrigerator light. Wet asphalt. Flickering ceiling tube. Steam from the kettle. Rain on one specific windowpane. A buzzing AC unit. Tight corridor walls. Mirror reflection. (See §9 — Environment plays.) 2. Physical micro-action on the body. The emotion translated into the actor's body. Jaw locks. Knuckles whiten. Lips press flat. Eyes drop a quarter-inch. He swallows hard. Fingers curl against the doorframe. The actor's body is the only place where feelings render — names of feelings do not. 3. Sound anchor or visual motif. A recurring perceptual hook tied to the spine of the piece. Stomach growl repeated three times. Reflection in dark glass on every transition. The same musical sting at every Crack beat. The clock's second hand. Footsteps in an empty corridor.
A shot with zero of these is filler. A shot with one is thin. Strong shots carry all three. The writer's job is to do this work — the model cannot infer it.
What is banned
Words that mark a writer being lazy. Each one is a placeholder for absent detail.
- "cinematic", "professional", "high quality", "masterpiece", "stunning", "epic", "amazing"
- "beautiful lighting", "dynamic camera", "intense moment", "powerful scene"
- "he is sad", "she is angry", "he is afraid" — emotions named without a body
Replace each with concrete physical facts. If a writer cannot do that, the scene is not yet thought through.
How details map to dramaturgy
Each layer of dramaturgy has a default detail register:
- Scene formula → environmental pressure (the geometry and atmosphere of the obstacle).
- Three-jobs rule → physical micro-action (the body shifts when emotion / action / pressure shifts).
- Five anchors → sound and visual motif (the motif is the anchor, the final image carries the motif).
If a shot's detail set does not match its dramaturgical function, the function will not land on screen.
3. The three-jobs rule
Every shot must do at least one of three things. If it does none, delete it.
- Change emotion (in hero, viewer, or the dynamic between characters).
- Advance action (new physical event, new information, new position).
- Increase pressure (stakes rise, clock ticks, space tightens, witness appears).
"Beautiful establishing shot" is not a job. "Beautiful hero shot of product" is not a job. Either the frame works for one of these three, or it is a fantik (wrapper without candy).
4. Walter Murch Rule of Six
From the editor of Apocalypse Now, The Godfather Part II, and The English Patient. The priority order when deciding where to cut. Each item is weighted heavier than the sum of everything below it.
1. Emotion (51%). Does the cut honor the emotional truth of the moment. What does the viewer feel now vs. what they should feel next. 2. Story (23%). Does the cut advance story or reveal character. 3. Rhythm (10%). Does the cut fall on a musical beat of the scene. 4. Eye-trace (7%). Where is the viewer's gaze at the moment of the cut. Does the new shot receive that gaze naturally. 5. 2D plane (5%). Does the cut respect the axis of screen direction. 6. 3D space (4%). Does the cut respect the geometry of the real location.
Practical consequence. Cutting for "динамика" (pace for its own sake) sits at item 3. If you cut there without serving items 1 and 2, the result is TikTok-ad for the attention-deficit.
5. Blocking as choreography of desire
Blocking is not "where the actor stands." Blocking is a visual answer to "what does the character want and from whom."
For every character in the scene, name.
- What they want now.
- Who or what they move toward.
- Who or what they move away from.
- Whom they corner.
- To whom they yield space.
- What gesture reveals the hidden desire.
Bad. "He stands near the window." Good. "He edges toward the window but his shoulder stays angled back toward her, as if the conversation still holds him."
6. Staging controls subtext
Staging is the arrangement of people, objects, and camera inside the frame. Before dialogue, staging already tells the conflict.
Power signals in staging.
- The standing character dominates the seated one.
- The character in the doorway controls the room.
- The character behind glass or in reflection is psychologically distant.
- The character in shadow carries threat or grief.
- Negative space around a character signals isolation.
- Tight framing signals suffocation.
- Shared frame without eye contact signals broken intimacy.
Spielberg, Kubrick, Iñárritu often build entire scenes where the staging states the conflict before a line is spoken.
Before writing a prompt, name the power dynamic the staging reveals.
7. Camera must have a reason
Fincher rule. Every camera movement answers "what changed?" If the answer is nothing, the camera is static.
Reasons for camera movement.
- A character made a decision and the camera follows the shift.
- New information arrived in the frame.
- Pressure escalated and the camera tightens.
- The character looked, and the camera reveals what they saw.
- A gesture pulled focus, and the camera rack-focused.
- The space changed (door opened, someone entered).
Bad. "Cinematic gliding camera movement." Good. "Push-in starts on 'I don't know' and stops on her jaw locking."
8. Spatial clarity beats montage hysteria
Spielberg principle. Even in chaos, the viewer must know.
- Where the hero is.
- Where the threat is.
- Which direction is escape.
- Which direction is decision.
High craft means fast, nervous, and still readable. Random whip-pans and strobe cuts without geography destroy drama. The fastest action scenes in cinema (Spielberg, Miller, early Bay) are built on a clear geometric map maintained through every cut.
Before writing a fast-cut sequence, sketch the geography in one sentence. "Hero moves left-to-right. Threat enters from the top of the frame. Exit is off-camera right."
9. Environment plays
Kurosawa principle. Weather and environment are characters. They amplify state.
Translate this to a short video by leaning on one environmental pressure.
- Flickering fluorescent light signals decay, bureaucracy, dread.
- Rain on a window signals grief withheld.
- Steam from a kettle signals suppressed anger.
- A buzzing air conditioner signals dissociation.
- Wet asphalt at night signals guilt.
- A tight corridor signals the walls closing in.
- A mirror or glass surface signals self-reckoning.
- Overhead cold office light signals judgment.
Pick one environmental pressure per scene and let it carry the emotion.
10. Three-layer storyboard method
Build storyboards in three layers, in this order. Skipping a layer produces pretty but empty output.
Layer 1. Dramatic beats
For a 60-90 second piece, a reliable beat map.
0-5s Hook. Hero already in tension. No setup. Problem on screen.
5-15s Context. Where we are. Who is near. What is at stake.
15-30s Pressure. Hero tries to hold control.
30-45s Crack. A detail appears that breaks the hero's position.
45-60s Acceleration. Cuts shorten. Breath tightens.
60-75s Impact. Decision, break, confession, or action.
75-90s Aftermath. Brief silence or visual residue.Adjust for 30s or 15s by compressing proportionally. Never skip the Crack or the Impact.
Layer 2. Shot functions
Tag every shot with a function. Same taxonomy as in role-modes.md, expanded with Power.
- Establish. Where we are.
- Power. Who controls the scene right now.
- Pressure. What is pushing down on the hero.
- Detail. Object, hand, phone, eye, drop, receipt, door. Macro anchor.
- Reaction. Face after the event.
- Shift. Inner change made visible.
- Impact. The decisive frame.
- Aftermath. Emptiness after action.
- Exit. Final image the viewer carries out.
This is cinema grammar. Everything else is decorative wallpaper.
Layer 3. Editing rhythm
Not random mincing. A rhythmic staircase.
long - shorter - shorter - pause - impactExample internal structure of an 8-10 second montage.
- 4s. Wide. Hero enters.
- 2s. Medium. Hero notices the object.
- 1s. Close-up. Eyes.
- 12 frames (0.5s). Macro insert of the object.
- 8 frames (0.33s). Hand.
- 6 frames (0.25s). Detail / sound cue.
- 2s. Sudden silence.
- 1s. Decision.
Rule. The pause before the impact is more important than the speed of the cuts. Without a pause, speed becomes visual meat grinder.
11. Shot card template
For each shot in a storyboard, fill every field. Missing fields reveal missing direction.
- Shot ID. 01, 02, 03
- Beat. What changes in the story here.
- Emotion. Fear, shame, anger, guilt, resolve, relief, etc.
- Frame. wide / medium / close-up / macro insert
- Composition. Center, edge, negative space, reflection, silhouette, foreground obstruction.
- Camera. Static, push-in, handheld, tracking, whip-pan.
- Movement reason. Why the camera moves here. Answer "what changed?"
- Action. Exact physical event.
- Eye trace. Where the viewer's gaze should land in the first 0.3s.
- Duration. 0.5s, 1s, 3s.
- Cut type. match cut, smash cut, cut on action, J-cut, L-cut.
- Sound. Breath, bass hit, street noise, phone ring, silence.
- Light / color. Cold, contrast, flicker, shadow, specific palette.
- Production note. Prop, location, actor direction.
If a shot card has empty fields, fill them or drop the shot.
12. Rhythm ladder
Drama rhythm is not uniform. It is stepped.
Slow-burn drama
4s, 4s, 3s, 2s, 1s, pause, 2s.
Commercial product arc
3s, 2s, 1.5s, 1s, 0.5s (product macro), 2s (hero shot).
Anxiety build
2s, 1s, 1s, 0.5s, 0.5s, 0.3s, pause, 1s.
Impact scene
pause, 0.2s (flash), 2s (aftermath in stillness).
Always insert at least one pause before the biggest cut.
13. One-anchor principle
For any short dramatic piece, commit to exactly five anchors. No more.
- One main emotion.
- One visual motif.
- One anchor object.
- One break.
- One final image.
Example.
- Emotion. Guilt.
- Motif. The hero keeps being reflected in glass surfaces (window, phone screen, elevator door).
- Anchor object. A phone with one unread message.
- Break. The hero deletes the message.
- Final image. The hero's face stays reflected in the darkened phone screen.
This set can be storyboarded, prompted, and cut together into something that carries real weight. It beats stacking "cinematic, professional, high quality, masterpiece" forever.
14. Worked example. Guilt, 30 seconds
Five anchors
- Emotion. Guilt.
- Motif. Reflections in glass.
- Anchor object. Phone with unread message.
- Break. He deletes it.
- Final image. His face ghosted on the dark phone screen.
Layer 1. Dramatic beats
- 0-3s. Hook. Phone buzzes on a dark desk. His face lit from below.
- 3-10s. Context. Small office after hours. Empty cubicles. Cold fluorescent. His reflection in the monitor.
- 10-18s. Pressure. He stares at the message preview. His jaw tightens. Hand hovers.
- 18-22s. Crack. He starts to type. Stops. Deletes character by character.
- 22-26s. Acceleration. Tight cuts. Finger. Screen. Eye. Breath. Window reflection.
- 26-28s. Impact. He taps Delete Conversation. One tap. Silence.
- 28-30s. Aftermath. Screen goes black. His face remains reflected in the dark glass.
Layer 2. Shot functions (selected)
- Shot 03. Establish. Empty office.
- Shot 05. Power. He is alone. The room looms.
- Shot 07. Detail. Phone screen close-up.
- Shot 09. Reaction. His face, jaw tightening.
- Shot 11. Shift. Hand hesitates over keyboard.
- Shot 13. Impact. Thumb taps Delete.
- Shot 14. Aftermath. Dark screen, his ghosted reflection.
Layer 3. Rhythm (final 10 seconds)
- 2s. Finger hovers over Delete.
- 1s. His face.
- 0.5s. Thumb.
- 0.5s. Screen confirmation prompt.
- 1s. His eyes.
- 0.5s. Thumb taps Delete.
- 3s silence. Screen goes black.
- 1.5s. His reflection on the dark phone.
Now each beat can be translated into a model-specific prompt using references/seedance.md, references/kling.md, or references/veo.md.
---
15. Dramaturgy check before sending any prompt
Run this six-point check before returning the final prompt to the user.
1. Is the scene formula complete? (desire + obstacle + geometry + gaze + rhythm) 2. Does every shot pass the three-detail check? (environmental pressure + physical micro-action + sound or visual motif) 3. Does every shot do one of the three jobs? (change emotion, advance action, increase pressure) 4. Is there a motivated reason for every camera move? 5. Is the spatial geometry readable? 6. Are the five anchors named? (emotion, motif, object, break, final image)
If any answer is no, fix before sending. Step 2 is the most violated.
Fixes, checklist, and cross-model skeletons
Contents
1. Continuity checklist (before final output) 2. Common failures and fixes 3. Cross-model prompt skeletons 4. Default negative constraints 5. Prompt compression order 6. Output format templates
---
1. Continuity checklist
Before sending a prompt to the user, verify.
- Same character across shots (face, body, hair).
- Same clothes across shots, exactly named.
- Consistent location logic.
- Object state progression makes sense (sausage in fridge -> on plate -> on fork -> on floor).
- No unwanted extra characters.
- No unwanted text or subtitles.
- No unwanted logos.
- Consistent age / body / face.
- No impossible hand-object action.
- No vague camera instructions ("cinematic camera").
- No vague lighting instructions ("beautiful light").
- No random stacking of director references.
- No contradictions.
- Palette is defined with concrete colors.
- Final image is explicitly stated.
If any item is missing, fix before sending.
---
2. Common failures and fixes
One continuous take instead of montage
Fix.
This must be a multi-shot sequence with visible hard cuts. Do not generate a single continuous take. Each beat uses a different angle and framing.For Seedance specifically, add explicit Cut to. or Camera cut to. markers in the prompt body.
Character face changes between shots
Fix.
Preserve the exact character in every shot. Same face shape, same eye color, same hair, same clothing, same expression style. [repeat the full identity block]For Kling. Use Element Binding with 3-4 reference images (front, side, three-quarter).
Object disappears mid-scene
Fix.
Track the object continuously. The same object remains visible or clearly implied in every beat.Describe the object's state progression in one sentence. "The sausage moves. fridge -> pot -> fork -> floor."
Weak drama. Scene feels flat.
Fix.
Play the scene with full emotional seriousness. Treat the ordinary object as if it carries life-or-death meaning. No comedy pacing. No detached observation.Messy random cuts in fast montage
Fix.
Use fast montage with clear readable action per cut. Every cut shows a distinct detail. face, hand, object, reaction, impact. Each cut must have a visible function.Dialogue too fast (Veo)
Fix. Cut the line. 8 seconds max of spoken text. Test by reading aloud at normal pace.
Melting hands / extra fingers
For Kling. Add to negative field.
distorted hands, extra fingers, melted face, deformedFor Seedance and Veo. Use positive phrasing.
anatomically correct hands, clean finger separation, realistic proportionsLighting drifts between clips
Fix. Name the dominant source and direction and repeat it verbatim in every clip.
Lighting constant. Cold fridge light as key from frame-right. Warm window spill as rim from frame-left. Same contrast ratio in every shot.Model ignores camera instruction
Fix. Move camera to the front of the prompt.
Bad. "A man opens a fridge. The camera is a slow push-in." Good. "Slow 50mm push-in. A man opens the fridge."
Weird AI-looking faces
Fix. Avoid style words like "hyperrealistic 8k masterpiece." These push the model into AI-art territory. Use production language instead.
Shot on 50mm, natural skin texture, motivated lighting, documentary feel.---
3. Cross-model prompt skeletons
Seedance
Subject. [identity block].
Motion. [one clear present-tense action].
Camera. Shot 1. [framing, lens, movement]. Cut to. Shot 2. [framing, lens, movement]. Cut to. Shot 3. [framing, lens, movement].
Environment. [location, time, props].
Lighting. [source, direction, quality, color].
Style. [realism level, genre reference].
Audio. [ambient, SFX]. (1.5+ only)
Continuity. [what must remain constant].
--resolution 1080p --duration 5 --camerafixed falseKling text-to-video
[Subject with identity anchor]. [Subject movement]. Scene. [3-5 environment elements]. Camera. [one movement + lens]. Lighting. [source + quality]. Atmosphere. [mood].
Negative field. blurry, distorted hands, extra fingers, melted face, watermark, subtitles, jitter.Kling image-to-video
Preserve [silhouette / feature / label]. [Camera movement, one lens]. [One or two motion verbs]. [Atmospheric cue, light change].
Negative field. blurry, distorted hands, melted face, jitter.Veo prose
[Subject performing action] in [environment]. [Camera framing + lens + movement]. [Lighting direction + color]. [Style + mood + palette].
Audio: [ambient, SFX, music texture].
Says: [character] says, "[dialogue, max 8s of speech]."
SFX: [punctual sound events].
Duration: [4 / 6 / 8] seconds.Veo JSON
See references/veo.md section 6 for full schema.
---
4. Default negative constraints
Where negatives are supported (Kling field, Veo body text, Seedance 2.0 fragile).
No subtitles. No on-screen text. No extra characters. No changing clothes. No changing face. No logos. No cartoon physics unless requested. No warm yellow tones unless requested. No random camera drift. No single-take when multi-shot is requested. No distorted hands. No extra fingers.For Kling field. Rewrite as positive entities. "distorted hands, extra fingers, subtitles, logos, cartoon physics, random camera drift."
For Seedance 1.0. Invert all negatives to positive phrasings in the prompt body.
---
5. Prompt compression order
When a model performs better with shorter prompts (Kling 2.5 Turbo, Kling 1.6, any image-to-video), cut in this order.
1. Keep character continuity. 2. Keep story action. 3. Keep shot timecodes (where relevant). 4. Keep lighting. 5. Keep camera. 6. Keep editing grammar. 7. Keep sound. 8. Remove philosophy and meta-commentary. 9. Remove extra adjectives. 10. Remove director references.
The goal. Preserve the skeleton. Lose the perfume.
---
6. Output format templates
Format A. Single prompt
One ready-to-copy prompt for one generation. Use the appropriate model skeleton.
Format B. Multi-clip prompts
Sequence of self-contained prompts. Each one repeats the full continuity block. Label them Clip 1 / 5, Clip 2 / 5, etc.
Between clips, add a one-line note explaining how they cut together. "Clip 1 ends on his hand reaching into the fridge. Clip 2 opens on his hand already inside the fridge, same light."
Format C. Storyboard (раскадровка)
Table with columns.
| Time | Shot | Function | Action | Camera | Light | Sound | Emotion |
|---|---|---|---|---|---|---|---|
| 0-1s | WS | Establish | Man walks to fridge | 35mm, slow push-in | Cold fluorescent overhead | Fridge hum | Exhaustion |
| 1-2s | MCU | Reveal | He opens fridge door | 50mm, static | Cold fridge light as key | Door seal pop | Anticipation |
Adjust row count to clip length.
Format D. Prompt audit
Given a user prompt. Return six sections.
1. What works. 2. What breaks generation. 3. Missing direction (camera, light, continuity). 4. Continuity risks. 5. Model-specific mismatches (wrong syntax for the chosen model). 6. Stronger version. Rewritten prompt, ready to copy.
Format E. Director treatment
For concept stage before any prompt is written.
- Core idea (one sentence)
- Emotional arc (three states)
- Visual motif (one recurring element)
- Rhythm (pace logic)
- Camera language (dominant grammar)
- Lighting (dominant source)
- Sound (texture)
- Ending image (final frame)
Format F. Veo JSON
Structured scene-by-scene JSON. Use for complex continuity. See references/veo.md section 6.
Kling reference (Kuaishou)
Contents
1. What Kling is 2. Versions and element limits 3. Kling 3.0 — multi-shot, native audio, 15s (read first if user is on 3.0) 4. Prompt formula (1.x – 2.x) 5. Prompt length by model 6. Negative prompts (dedicated field, special rule) 7. Element Library and Element Binding (unique to 1.x – 2.x) 8. Motion Brush (unique) 9. Motion Control (2.6 Pro, unique) 10. Image-to-video rule 11. Failure modes and fixes 12. Skeleton and example
---
1. What Kling is
Strong at realistic physics, character animation, and consistency through reference images. Fast generation. Best for social-ready clips and repeat-character scenes.
Kling 3.0 changed the model's positioning fundamentally — it now competes with Seedance on multi-shot output (up to 6 shots per generation) and with Veo on native dialogue and lip-sync. See section 3.
2. Versions and element limits
Kling models vary widely in how many distinct elements they can handle in a single prompt. Overstuffing produces melting faces, broken hands, or static motion.
- Kling 1.6. Simplified prompts. Keep it very simple.
- Kling 2.1 Pro. 1080p at 30fps. Motion Brush. End frame control.
- Kling 2.5 Turbo Pro. Maximum 3-4 distinct elements.
- Kling 2.6 Pro. 5-7 elements. Motion Control. Element Binding for character consistency.
- Kling 3.0. Multi-shot in one generation (up to 6 shots). Native audio with dialogue and lip-sync. 15s continuous output. Strongest character and scene consistency. Two tiers: v3/pro and v3/standard.
If in doubt about the user's version, ask. The prompting protocol differs between 1.x – 2.x (sections 4-9) and 3.0 (section 3).
3. Kling 3.0 — multi-shot, native audio, 15s
Kling 3.0 is a different beast. It understands cinematic intent, not just visual descriptions. Prompts read like scene directions, not object lists.
What changed vs 2.x
| Capability | 1.x – 2.6 | 3.0 |
|---|---|---|
| Multi-shot in one generation | No (one continuous take) | Yes — up to 6 shots |
| Native audio | Limited / off | Yes — dialogue, ambient, voice tone |
| Lip-sync | No | Yes — coherent across multi-character scenes |
| Max duration | 5-10s | Up to 15s |
| Character labeling | Via Element Library + reference images | Via in-prompt `[Character A: ...]` tags (still supports references) |
| Cinematic language understanding | Partial | Full — reads "shot-reverse-shot", "POV", "tracking shot", "macro close-up" |
Five-layer prompt structure
Scene → Characters → Action → Camera → AudioWrite each layer as flowing prose with explicit labels.
Anchor subjects early
Introduce every character and key object at the start of the prompt, before any shot description. Use unique consistent identifiers — the same label survives across all shots:
[Character A: Exhausted Partner — late 40s, gray-streaked beard, navy peacoat, hollow eyes]
[Character B: Female Investor — early 30s, sharp blazer, calm posture]Then refer to them as "Character A" / "Character B" inside shot descriptions. This locks identity better than re-describing the character each shot.
Multi-shot syntax (think in shots, not clips)
[Character A: ...]
[Character B: ...]
Master intent: tense negotiation in a glass-walled office at dusk.
Shot 1 (0-3s). Wide tracking shot. Character A enters frame from left, crosses to the table.
Shot 2 (3-6s). Profile close-up on Character A. He sets down a brown leather folder.
Shot 3 (6-9s). Shot-reverse-shot. Cut to Character B's face. She does not blink.
Shot 4 (9-12s). Macro insert on her hand tightening around a fountain pen.
Shot 5 (12-15s). Two-shot, low angle, both reflected in the glass wall behind them.
Camera. Slow, deliberate. Sony FX6 feel, 35mm and 85mm.
Lighting. Cold blue dusk through floor-to-ceiling windows. Single amber desk lamp.
Audio. Distant city ambience. No music. Footsteps on hardwood. Pen clicks on paper.Each shot must answer: framing + subject + motion. Empty shot descriptions ("static frame, ambient mood") collapse into one continuous take.
Dialogue protocol (P1-P4)
For any speaking scene, follow these four rules. The fal.ai guide calls them P1, P2, P3, P4.
P1. Structured naming. Use unique identifiers per character.
- ✓
[Character A: Black-suited Agent]and[Character B: Female Assistant] - ✗ "[Agent] says... Then, he says..."
P2. Visual anchoring before dialogue. Bind dialogue to a unique action first.
- ✓ "Character A pulls a folded note from his pocket and reads aloud: 'It's not what you think.'"
- ✗ "Character A says 'It's not what you think.'" (no visual anchor → lip-sync drifts)
P3. Voice tone in the tag. Assign emotion / texture inline.
- ✓
[Character A, raspy deep voice]: "We're out of time." - ✓
[Character B, clear fearful voice]: "Don't open it." - ✗
[Man]: "We're out of time."(vague — model picks a generic voice)
P4. Temporal control between lines. Use linking words to prevent dialogue from merging.
- ✓ "Character A: 'I won't ask again.' Immediately, Character B: 'You don't have to.'"
- ✗ Two consecutive
[Character X]: "..."lines with no transition (model overlaps them)
Image-to-video on 3.0 — lock first, then move
The input image serves as the anchor for identity, layout, and on-image text. Keep the prompt short and motion-focused. Describe how the scene evolves from the image, not what is in the image.
Preserve identity, wardrobe, and the storefront sign exactly.
[Character A: same person from the image]
Shot 1 (0-2s). She turns her head toward camera, exhales.
Shot 2 (2-5s). Slow push-in to a tight close-up. Wind catches her hair.
Audio. Distant traffic, wind through awnings, a soft bell from inside the shop.Cinematic vocabulary that 3.0 actually understands
Use these terms directly — the model treats them as instructions, not flavor:
- Framing: profile shot, three-quarter, macro insert, two-shot, OTS, POV, low angle, high angle, Dutch
- Edits: shot-reverse-shot, match cut, smash cut, J-cut, L-cut
- Camera moves: tracking, dolly, push-in, pull-out, whip pan, crane, handheld
- Lens feel: 35mm, 50mm, 85mm, 100mm macro, anamorphic 40mm
Default model choice
If the task hits any of these — use Kling 3.0 over earlier Kling:
- Multi-shot dialogue scenes
- 10-15s continuous narrative
- Lip-sync required
- Multi-character with distinct voices
- Image-to-video where text on the image must stay legible
For everything else (single clip, no dialogue, < 10s, character lock via reference images, Motion Brush) — older versions still work and are cheaper.
4. Prompt formula (1.x – 2.x)
For 1.6, 2.1 Pro, 2.5 Turbo Pro, 2.6 Pro:
Subject (with specific details)
+ Subject Movement (one clean verb phrase)
+ Scene (3-5 elements max)
+ Camera Language
+ Lighting
+ AtmosphereWrite as flowing prose. Kling 1.x – 2.x dislikes fragmented tag-style inputs.
For Kling 3.0, use the multi-shot structure from section 3 instead.
5. Prompt length by model
- 1.6. Keep it simple, minimal.
- 2.5 Turbo Pro. 3-4 elements max. 50-80 words.
- 2.6 Pro. 5-7 elements. 50-80 words.
- 3.0. Longer prompts welcome — multi-shot needs explicit structure (see section 3). Plan ~30-60 words per shot.
- Image-to-video (any version). 20-40 words. Shorter, motion-focused.
Long prompts on 1.x – 2.x = melted outputs. Compress ruthlessly. On 3.0, structure beats length.
6. Negative prompts (critical rule)
Kling has a dedicated negative prompts field. The field auto-interprets input as exclusion. Do not write "no X". Write the thing itself.
Bad. "no robots" Good. "robots"
Bad. "no blurry faces, no distorted hands" Good. "blurry faces, distorted hands, melted features"
Common effective negatives.
low resolution, blurry, distorted hands, extra fingers, melted face, watermark, subtitles, logo, text overlay, jitter, shaking camera, deformedKeep the list short. Long negative stacks reduce motion and detail.
7. Element Library and Element Binding (1.x – 2.x)
For character consistency across generations, Kling has an Element Library. Upload 3-4 reference images of the character from different angles.
Required angles.
- Front
- Side (profile)
- Three-quarter
Then in Image-to-Video settings, enable "Bind Elements" to lock features. This gives the AI a visual anchor that survives camera pans and light changes.
Source image rules.
- 1080p or higher.
- Even lighting. Avoid hard shadows. The AI can mistake them for permanent facial features.
- No text or watermarks.
- Clean uncluttered background.
- Centered subject.
- Well-contrasted.
For Kling 3.0 — Element Library still works, but the in-prompt [Character A: ...] labels (see section 3) are usually enough on their own.
8. Motion Brush (unique)
Animate up to 6 regions of a single image independently. Each region gets its own motion path.
Critical rule. The text prompt MUST match the brush motion. If you brush a river flowing and write "stagnant pond" the model tears itself apart. Align prompt verbs with brush directions.
9. Motion Control (2.6 Pro, unique)
Copy motion from a reference video. Use when you need specific performance or choreography (dance, martial arts, specific gait). The prompt then focuses on subject description only, the reference handles motion.
10. Image-to-video rule
Keep the prompt short (20-40 words). Focus only on motion. Do not re-describe static elements the model already sees.
Include explicit continuity cues. Kling responds well to "preserve X" instructions.
Example.
Preserve silhouette and label text. Slow tracking shot from the side. She turns her head toward camera. Wind catches her hair. 35mm, golden hour rim light.For Kling 3.0 image-to-video — see section 3 ("Image-to-video on 3.0 — lock first, then move").
11. Failure modes and fixes
Random camera drift in static scenes
Fix. Say "locked static frame" or "camera fixed, no movement."
Prompt exceeds model capacity, output melts
Fix. Cut to model-appropriate element count. 3-4 for Turbo, 5-7 for 2.6 Pro.
Character face changes between generations
Fix. Upload 3-4 reference images to Element Library. Use Bind Elements in the settings.
Negative prompts ignored
Fix. Rewrite "no X" as "X" in the negative field.
Motion Brush artifact
Fix. Check that the text prompt verbs match the brush direction. Rewrite the text if it contradicts the motion.
Element overload melts hands
Fix. Simplify. Combine elements where possible. Cut secondary descriptions.
Kling 3.0 collapses multi-shot into one continuous take
Fix. Make shot boundaries explicit: Shot 1 (0-3s). ... Shot 2 (3-6s). .... Each shot description must include a different framing or camera angle. If two adjacent shots share both framing and angle, the model merges them.
Kling 3.0 dialogue lip-sync drifts
Fix. Apply P2 — bind dialogue to a unique visual action ("Character A pulls a folded note from his pocket and reads aloud:") before the line itself. Then add P4 linking words ("Immediately,") between consecutive lines.
Kling 3.0 picks the wrong voice
Fix. Apply P3 — put voice tone inside the speaker tag: [Character A, raspy deep voice]:. Generic [Man]: / [Woman]: tags get a generic voice.
12. Skeleton
Text-to-video (1.x – 2.x)
[Subject with identity anchor]. [Subject movement in one clean verb phrase]. Scene. [3-5 environment elements]. Camera. [one movement + lens]. Lighting. [source + quality]. Atmosphere. [mood].
Negative field. blurry, distorted hands, extra fingers, melted face, watermark, subtitles, jitter.Image-to-video
Preserve [silhouette / label / specific feature]. [Camera movement, one lens]. [One or two motion verbs]. [Atmospheric cue, light change].
Negative field. blurry, distorted hands, melted face, jitter.Worked example. Fashion, image-to-video (2.6 Pro)
Preserve silhouette and fabric texture. Slow lateral tracking, 85mm. She turns her head toward camera, exhales through her nose, hair catches golden hour wind. Warm rim light. Confident stillness.
Negative field. blurry, distorted hands, extra fingers, melted face, watermark, motion blur, jitter.Multi-shot with dialogue (Kling 3.0)
[Character A: Investigator — late 30s, navy raincoat, tired blue eyes, three-day stubble]
[Character B: Witness — early 20s, oversized hoodie, hands wrapped around a paper cup, eyes red from crying]
Master intent: a 12-second interrogation in a fluorescent-lit precinct break room at 2am. The investigator holds back, the witness breaks.
Shot 1 (0-3s). Two-shot, eye level, 35mm. Character A sits across from Character B. Static frame.
Shot 2 (3-6s). OTS over Character A's shoulder. Character B's hands tighten on the paper cup.
Shot 3 (6-9s). Profile close-up on Character A. He slides a photo across the table.
[Character A, low even voice]: "You were there."
Shot 4 (9-12s). Reverse — close-up on Character B. She does not look up.
Immediately, [Character B, fragile broken voice]: "I didn't see anything."
Camera. Locked frames, no movement except the slide of the photo. Sony FX6 feel.
Lighting. Cold overhead fluorescent, slight flicker. Pale skin, hard shadows under the eyes.
Audio. Distant police radio chatter, the buzz of the fluorescent tube, a vending machine humming in the corridor. No music.
Negative field. blurry, distorted hands, extra fingers, melted face, watermark, subtitles, jitter, dialogue overlap.Montage patterns and genre modules
Contents
1. Montage patterns (6 ready structures) 2. Genre modules (7 archetypes) 3. Multi-clip story structure
---
1. Montage patterns
These are pre-built structures that solve common scene types. Pick one, fill with your specifics.
Pattern 1. Escalation
Use for tension builds, reveals, dramatic emphasis.
wide -> medium -> close-up -> macro -> reaction close-up -> impactPattern 2. Anxiety
Use for psychological pressure, internal conflict, impending bad news.
face -> object -> hand -> face -> object closer -> sound cue -> sudden stillnessPattern 3. Discovery
Use for revealing a hidden element, exploration, search.
POV -> empty space -> searching hand -> hidden object -> rack focus -> emotional reactionPattern 4. Catastrophe
Use for comedic or dramatic disaster. Tiny object failures play bigger than explosions.
anticipation -> object instability -> reaction -> object falling -> impact -> silence -> emotional collapsePattern 5. Commercial product drama
Use for ads with a hero product.
lifestyle setup -> product reveal -> macro texture -> human reaction -> use moment -> hero product shotPattern 6. Music video loop
Use for rhythmic repetition, transformation, performance.
gesture A -> cut -> gesture A from new angle -> cut -> transformation -> repeat gesture A -> release---
2. Genre modules
Each genre has a distinct visual grammar. Match style and rhythm to the genre.
Domestic tragedy
Ordinary objects treated with extreme seriousness. Tiny event as cosmic disaster.
- Style. Cold night interior. Mundane location. Dramatic close-ups. Slow push-ins. Macro inserts. Sudden silence.
- Tone. Tragicomic. Absurd sincerity.
- Pace. Slow build, longer reaction shots.
- Example. A man discovering his fridge is empty, played like a Greek tragedy.
Music video
Rhythm. Repetition. Visual motif. Transformation.
- Style. Repeated gestures. Match cuts. Performance fragments. Visual loops. Color-coded sections. Aggressive lens changes.
- Tone. Emotional intensity. Sensory overload with readable structure.
- Pace. Fast, beat-driven.
Commercial
Clarity. Product logic. Sensory detail.
- Style. Clean lighting. Intentional macro. Clear product visibility. Controlled movement. Readable final hero shot.
- Tone. Desire. Transformation. Benefit shown through action.
- Pace. Medium, building to hero shot.
Psychological drama
Pressure. Stillness. Negative space.
- Style. Locked frames. Long close-ups. Reflections. Obstructed framing. Quiet sound design.
- Tone. Internal conflict. Hidden tension. Emotional compression.
- Pace. Slow, sustained.
Action
Spatial clarity above all else.
- Style. Establishing geography. Clear direction of movement. Impact inserts. Wide shots between close-ups. Strong eyeline continuity.
- Tone. Urgency. Force. Readable chaos.
- Pace. Fast but geometrically clear.
Fashion
Symmetric framing. Controlled palette. Repeated silhouettes.
- Style. Slow motion micro-beats. Macro fabric detail. Confident stillness. Rim light.
- Tone. Assured. Sensual.
- Pace. Slow, intentional.
UGC / Social
Authentic. Vertical. Quick.
- Style. Handheld. Natural light. 9:16. Vertical compositions. Direct-to-camera gestures. Quick beats.
- Tone. Immediate. Casual. Urgent relevance.
- Pace. Rapid, thumbnail-readable.
---
3. Multi-clip story structure
AI video generators have no memory between generations. Longer videos live in multiple clips stitched in the edit.
Default splits
- 10s. 2 clips x 5s
- 15s. 3 clips x 5s
- 30s. 6 clips x 5s
- 45s. 9 clips x 5s
- 60s. 12 clips x 5s
Exception. Seedance multi-shot can pack 2-3 shots into a single 5-10s clip.
What every clip must repeat
Every clip is self-contained. The model has no memory. Repeat this block in every clip.
- Character identity (face, hair, clothing, distinguishing marks)
- Clothing items, exactly named
- Location description
- Visual style (palette, grade, reference)
- Camera language (dominant grammar of the whole piece)
- Lighting logic (dominant source and direction)
- Continuity rules (what must remain constant)
- Color palette with specific colors
Yes, every single clip. Yes, even if it feels repetitive. That repetition is the only thing keeping the output consistent.
Timecoded internal beats
Inside each 5-second prompt, use timecodes.
0.0-0.8. [action / camera / light]
0.8-1.6. [action / camera / light]
1.6-2.5. [action / camera / light]
2.5-3.7. [action / camera / light]
3.7-5.0. [action / camera / light]Density by genre.
- Emotional drama. 3-4 beats per 5s. Longer reactions.
- Standard narrative. 4-7 beats per 5s.
- Fast montage / music video. 6-9 beats per 5s.
Use timecoded structure for Seedance and Veo. Kling prefers flowing prose.
Continuity across clips
When moving from clip N to clip N+1, the first beat of the new clip should match the final beat of the previous clip. Character in same pose. Same light. Same color temperature. This is what makes them cut together cleanly in the editor.
Example. Clip 1 ends on him reaching into the fridge. Clip 2 opens on his hand inside the fridge at the same height and light.
Role modes
This skill fuses three mindsets. Switch role based on the task stage. The dramaturgy layer sits underneath all three roles. For the full dramaturgy method (scene formula, Murch Rule of Six, blocking as choreography of desire, staging and subtext, three-layer storyboard, shot card template, rhythm ladder), load references/dramaturgy.md. Director mode leans heaviest on that file.
Contents
1. Director mode 2. Screenwriter mode 3. Editor mode 4. Shot function taxonomy
---
1. Director mode
Trigger
User asks to "придумать сцену", "разработать концепцию", "визуализировать идею", "сделай как Финчер", "предложи трактовку", "какой вайб", or shares a raw idea without structure.
Mindset
Every shot must answer at least one of these questions. If a shot answers none, delete it.
- What changes emotionally?
- What new information appears?
- What action moves the story?
- What pressure increases?
- What does the viewer need to notice?
- Why does the camera move?
- Why does the cut happen here?
Output format. Director treatment
Core idea. [one sentence, the point of the piece]
Emotional arc. [start -> middle -> end, three emotion states]
Visual motif. [one recurring visual element]
Rhythm. [pace logic. slow build, rapid montage, still + sudden burst]
Camera language. [dominant grammar. handheld intimacy, locked precision, gliding observer]
Lighting. [dominant source and direction]
Sound. [texture. ambient layers, music role, silence moments]
Ending image. [final frame in one sentence]Style references
Use one dominant director reference per scene. Translate into concrete camera, light, rhythm, blocking. Never stack three or more references.
- Fincher. Precise motivated camera. No handheld drift. Cold palette.
- Kurosawa. Weather as emotional pressure. Rain, wind, heat.
- Spielberg. Readable staging. Clear geography. Emotional wide shots.
- Edgar Wright. Sound-driven montage. Match cuts on sound events.
- Jonathan Glazer. Music-video visual idea translated to drama.
- Wong Kar-wai. Longing. Reflections. Slow emotional drift.
- Safdie. Anxiety. Handheld pressure. Overlapping sound.
- Wes Anderson. Symmetry. Controlled blocking. Color blocks.
---
2. Screenwriter mode
Trigger
User asks to "напиши сценарий", "разбей на биты", "придумай диалог", "адаптируй историю в видео", "нужен story arc."
Mindset
Translate every beat into physical action the camera can see. Tag each beat with a shot function (see section 4). Every beat must have subtext.
Output format. Beat breakdown
Beat 1. [function] - [physical action]. Subtext. [what it means].
Beat 2. [function] - [physical action]. Subtext. [what it means].
...
Final image. [what the viewer takes with them].Dialogue format (for Veo specifically)
Write spoken lines in ready-to-paste Veo syntax.
He says, in a flat exhausted voice, "We are fine. We are fine."Cut lines until they fit 8 seconds of natural speech.
Internal monologue rule
Veo cannot render internal monologue. Translate into visible bodily signal.
Bad. "He is thinking about his father." Good. "He stops mid-motion. His eyes drift off-axis. He swallows. He resumes."
---
3. Editor mode
Trigger
User asks to "собери монтаж", "задай ритм", "сделай динамично", "разбей на склейки", "как смонтировать", "сколько кадров в 5 секундах."
Mindset
Dynamic montage is built through structure, not speed. Random fast cuts create visual soup. Readable fast cuts create propulsion.
Rules
- Clear shot function per cut.
- Escalating frame tightness (wide -> medium -> close).
- Different camera angles per cut.
- Motivated cuts (sound event, action beat, emotional shift).
- Contrast speed and pause.
- Sound-driven transitions.
- One silence or stillness moment before a major impact.
Output format. Timecoded beat sheet
00.0-00.8. [shot function] - [framing] - [camera] - [sound cue]
00.8-01.6. [shot function] - [framing] - [camera] - [sound cue]
...Density
- Emotional drama. 3-4 beats per 5s.
- Standard narrative. 4-7 beats per 5s.
- Fast montage. 6-9 beats per 5s.
More than 9 beats in 5s creates incoherent motion regardless of model.
---
4. Shot function taxonomy
Every beat should carry at least one function tag. This is the editor's grammar.
- Establish. Where we are. Wide or master shot. Sets geography.
- Reveal. What appears. New information enters frame.
- Power. Who controls the scene right now. Staging reveals the hierarchy before dialogue does.
- Pressure. Tension builds. Movement toward threat or decision. Environment can carry this (flickering light, rain, steam, a tight corridor).
- Detail. Important object, hand, eye, texture, gesture. Macro or close-up. The anchor object often lives here.
- Reaction. Emotional consequence. Face response to a beat.
- Shift. Inner change. Body language turning point. The moment before the decision becomes visible.
- Impact. Decisive visual event. The drop. The hit. The break.
- Aftermath. Emotional residue. Stillness after impact.
- Exit. Final state. The image the viewer leaves with.
A full scene usually moves. Establish -> Power -> Pressure -> Detail -> Reaction -> Shift -> Impact -> Aftermath -> Exit.
A montage sequence can compress or loop. Use function tags to keep the structure readable even at high speed.
Each tag is also a question the shot must answer. Establish. Where. Power. Who commands. Pressure. What pushes. Detail. What to notice. Reaction. What it cost. Shift. What changed inside. Impact. The moment. Aftermath. The residue. Exit. The image carried out.
Seedance reference (ByteDance)
Contents
1. What Seedance is 2. Versions and specs 3. CLI parameters 4. The Details Law (read first) 5. 6-step prompt formula (quick) 6. Production-grade skeleton (11 blocks) — for dramatic / multi-shot work 7. The 5-second shot timeline (rhythm template) 8. Multi-shot syntax (unique capability) 9. Anti-mush guard block (when Seedance smears the cuts) 10. @img1 character reference syntax 11. Camera movements (9 presets) 12. Negative prompts handling 13. Image-to-video rule 14. Audio (1.5+) 15. Failure modes and fixes 16. Worked example. 15-second tragicomedy as 3 × 5s clips
---
1. What Seedance is
A ByteDance video model with the UNIQUE ability to generate several distinct shots in ONE generation. The only major model that can pack a mini-montage inside a single 5-10s clip. Good cinematic motion, strong at ads and short narrative.
2. Versions and specs
- Seedance 1.0 Pro. 1080p, 5-10s (API up to 12s). Strong multi-shot.
- Seedance 1.0 Lite. 720p, faster, cheaper.
- Seedance 1.5 Pro. Native audio and lip-sync added.
- Seedance 2.0. Improved motion, better audio, 9 camera movement presets, limited negative prompt support.
Resolutions. 480p, 720p, 1080p. Aspect ratios. 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, 9:21. Frame rate. 24-30 fps.
3. CLI parameters
Append at the end of the prompt.
--resolution 1080p --duration 5 --camerafixed false --seed 42--resolution. 480p | 720p | 1080p--duration. 2 to 12 (Pro)--camerafixed. true locks the camera. false allows movement.--seed. for reproducibility
4. The Details Law (read this first)
Details intensify emotion. Laziness kills the prompt.
Seedance does not render abstractions. It renders physical specifics. Every adjective must be a sensory fact. Every emotion must be a body. Every shot must own at least three concrete details:
1. One environmental pressure. Cold blue refrigerator light. Steam off boiling water. Wet asphalt. Flickering fluorescent. Dripping tap. Curtain breathing in the AC. 2. One physical micro-action. Jaw locks. Finger taps the counter. Knuckles whiten on the fork. Lips press into a line. He swallows hard. 3. One sound anchor or visual motif. Stomach growl at 2.3s. Reflection in the dark phone screen. Rain hitting the same windowpane.
If a shot has none of these — it is filler. Delete it or rewrite it.
Banned, lazy phrasing that produces mush:
- "cinematic, professional, high quality, masterpiece"
- "beautiful lighting"
- "epic scene"
- "amazing visuals"
- "he is sad / he is angry" (with no physical translation)
This rule is hard. Multi-shot prompts on Seedance fail not because of the model, but because the writer was lazy on a single shot. One thin shot drags the whole sequence down.
5. Six-step prompt formula (quick scenes)
For single-shot 5s clips with one clear action.
Subject. [Who or what]
Motion. [Present-tense verb, one clear action]
Camera. [Movement + framing + lens]
Environment. [Where, props, atmosphere]
Lighting. [Direction + quality + color]
Style. [Realism level, genre reference]Write in full sentences, not tags. Seedance prefers clear grammatical prose. For dramatic / multi-shot / character-locked work, use the production-grade skeleton in section 6 instead.
6. Production-grade skeleton (11 blocks)
Use this for any dramatic piece, multi-shot ad, music-video segment, or character-locked clip. Each block answers a specific failure mode. Skipping a block reintroduces that failure.
[1] Character lock.
Use @img1 as the main character reference and preserve the exact same person
across the whole clip: <face shape, eye color, hair, facial hair, build, distinctive
features>. Dress him in <exact wardrobe>. No nudity. No glasses (unless reference).
No extra characters.
[2] Length + genre + editing intent.
Generate a <duration>-second multi-shot <genre> sequence with <fast / slow / staircase>
dynamic editing.
[3] Story (one paragraph).
<Concrete physical events of this clip in present tense. What changes from start to
end. Name the break point.>
[4] Visual style.
<Palette, contrast, grain, color temperature, what to avoid (e.g. "no warm yellow
tones"). Texture cues — realistic skin, food detail, fabric, surfaces.>
[5] Camera style.
<Camera body / look (e.g. Sony FX3 handheld). Lenses with purpose:
35mm / 50mm for medium, 85mm for emotional close-ups, 100mm macro for inserts,
24mm for wide silhouettes. Handheld micro-shake or static. Strict cuts between
shots. No continuous take.>
[6] Editing style.
<Hard cuts vs match cuts. Rhythmic escalation. Where the pause lands. Where the
impact hits. The rhythmic staircase: long → shorter → shorter → pause → impact.>
[7] Audio.
<Diegetic sounds in order: ambient, micro-actions, impact, silence moment, final
cue. Examples: stomach growl, refrigerator hum, fork clink, wet thud, abrupt
silence, distant city ambience. No dialogue / No subtitles / No on-screen text.>
[8] Shot-by-shot timeline.
Shot 1, 0.0–X.X sec: <framing, lens, camera move, action, environment detail,
emotion translated into body>.
Shot 2, X.X–Y.Y sec: <...>.
...
(See section 7 for the 5-second rhythm template.)
[9] Lighting (recap and specifics).
<Main source, fill, rim. Direction. Color temperature. What it carries
psychologically — judgment, isolation, hope, grief.>
[10] Composition.
<Where the subject sits in frame across shots. Negative space. Reflections.
Silhouettes. Foreground obstruction. The final image must be named.>
[11] Output specs.
<Exact duration. Aspect ratio. Realism level. CLI: --resolution 1080p
--duration 5 --camerafixed false>Why 11 blocks and not 6? Each block prevents one specific Seedance failure: identity drift, one-take collapse, mood smear, lens chaos, audio mismatch, drift on rhythm. Cheap insurance.
7. The 5-second shot timeline (rhythm template)
For a 5-second multi-shot clip, the model performs best with 5 shot beats following a dramatic micro-arc. Use this timing as the default scaffold:
0.0–0.8 sec | Establish | extreme close-up or insert that anchors emotion / situation
0.8–1.6 sec | Action | medium shot, hero moves or reacts
1.6–2.5 sec | Turn | new framing reveals the shift (POV, OTS, rack focus)
2.5–3.6 sec | Reaction | tight close-up, slow push-in, emotion lands on the body
3.6–5.0 sec | Climax / hero | hero shot, low angle, slow-mo if earned, final imageFor 10s clips, double the structure or insert one pause before the climax (the pause beats speed). For 15s+ stories, split into 3 × 5s clips and stitch in the editor — Seedance is more reliable in shorter generations than one long take.
8. Multi-shot syntax (unique)
Seedance reads explicit cut markers inside a single prompt and generates distinct shots connected by visible cuts. This is its strongest card. Use it when you need montage in a single generation.
Supported cut markers.
Shot 1. [description]
Cut to. [description]
Camera cut to. [description]
Camera switching. [description]
Lens switch to. [description]Inline timeline syntax.
[Shot A description] -> Cut to -> [Shot B description] -> Camera cut to -> [Shot C description]Example.
Shot 1. Medium close-up on a tired man in a kitchen. He opens the refrigerator.
Cut to. Macro insert. His hand reaches toward a single sausage on an empty shelf.
Cut to. Over-the-shoulder shot. The fridge light paints his face cold blue.Use 2-3 shots per 5-second clip for tight cinematic montage, 4-5 shot beats only when each beat is short and physically distinct (see section 7). Avoid 6+ shots per 5s. They compress into incoherent motion.
9. Anti-mush guard block
Seedance sometimes ignores cuts and produces one continuous take, smears multiple shots into a single moving frame, or drifts character identity across the timeline. Drop this block at the very top of the prompt (before block [1]) when it happens or to prevent it pre-emptively on heavy multi-shot work:
Important direction:
This must be a clearly edited multi-shot sequence with visible cuts between shots.
Do not generate a single continuous take. Each shot must have a different camera angle
and different framing. Use rapid montage pacing with rhythmic escalation. Keep the same
character appearance throughout. Preserve the same clothing, face, body type, facial
hair, and hairstyle in every shot. The tone is <serious cinematic drama / tragicomedy /
documentary realism / etc.>.This is the highest-leverage paragraph in any Seedance prompt. Add it any time the model produced mush on a previous attempt.
10. @img1 character reference syntax
Seedance 2.0 accepts image references inline using @img1, @img2, etc. Use this to lock the protagonist's likeness across all shots in a multi-shot clip and across multiple stitched clips.
Use @img1 as the main character reference and preserve the exact same man across the
whole clip: <full identity block: face, eyes, hair, facial hair, build, distinctive
features, wardrobe>. No nudity. No glasses (unless in reference). No extra characters.Rules.
- The full identity block must follow the
@img1mention. The model needs the textual description as a backup signal — the image alone drifts. - Repeat the identity block in every clip of a stitched sequence. Treat each generation as briefing a new intern.
- For multiple references (character + setting, character + outfit), label each:
@img1is the protagonist,@img2is the location reference,@img3is the outfit reference. Then state the role inline.
11. Camera movements (9 presets)
- Dolly Out. Reveals context, pulling away.
- Dolly In. Pushes into subject, builds tension.
- Pan Left / Pan Right. Horizontal reveal, landscape, row.
- Tilt Up / Tilt Down. Vertical reveal.
- Tracking. Follows a moving subject.
- Crane / Aerial. Epic scale, establishing.
- Handheld. Documentary, intimate, UGC.
- Zoom In / Zoom Out. Tension or detail.
- Hitchcock Zoom. "dolly out while zooming in" for vertigo effect.
- Static (via
--camerafixed true). Locked frame.
Combine sparingly. Multiple camera moves in 5s rarely resolve cleanly.
12. Negative prompts
Seedance 1.0 Pro does NOT support negative prompts. No --no blur syntax works.
Seedance 2.0 adds limited negative prompt support but it's fragile and often ignored.
Workaround. Always invert to positive phrasing. Instead of "no yellow tones" write "cold blue-gray palette with desaturated skin tones." Instead of "no distorted hands" write "anatomically correct hands with clear finger separation."
13. Image-to-video rule
When using a reference image, DO NOT describe elements already visible in the image. The model sees it. Describe only motion and camera work. Re-describing static elements creates identity drift.
Bad. "A man in red shirt stands in kitchen. He walks to the fridge." Good. "He slowly walks toward the fridge, opens it with hesitation, freezes when he sees the empty shelves. Tracking shot from behind, 35mm."
14. Audio (1.5+)
Seedance 1.5 Pro and 2.0 generate native audio and support lip-sync. Include audio cues in the body of the prompt.
Audio. fridge hum, distant rain on window, one stomach growl at 2.3 sec, final silence.Dialogue is supported but less robust than Veo. Keep lines short.
15. Failure modes and fixes
One continuous take when multi-shot was requested
Fix. Add explicit "Cut to" markers. Say "multi-shot sequence with visible hard cuts. Do not generate a single continuous take."
Character drift across shots
Fix. Repeat the full identity block at each shot boundary inside the prompt.
Motion ignored below 5 seconds
Fix. Minimum duration 5s for any scene that has multiple actions or camera moves.
Negative phrasing ignored
Fix. Use positive substitutes. Seedance 1.0 has no negative parser.
Multi-shot fails on 4+ shots in 5s
Fix. Cap at 2-3 shots per 5-second clip for tight cinematic, or 4-5 short beats following the section 7 timeline. If more shots are needed, split into multiple generations.
Mood smear / "everything looks the same"
Fix. Each shot needs a distinct emotional function (Establish / Power / Pressure / Detail / Reaction / Shift / Impact / Aftermath — see dramaturgy.md §10, Layer 2). If two adjacent shots have the same function, the model averages them into the same frame. Vary function and framing in every cut.
Lazy abstract phrasing produces lifeless clips
Fix. Apply the Details Law (section 4). Audit your draft: every shot must have one environmental pressure, one micro-action, one sound or visual motif anchor. Replace adjectives like "dramatic", "intense", "beautiful" with concrete physical facts.
16. Worked example. 15-second tragicomedy as 3 × 5s clips
A 15s narrative is never one prompt. It is three self-contained 5-second prompts, each with the full character lock, the full visual style, the full audio block, and a different dramatic function. Stitch in the editor.
Story spine
- Beat 1 (Clip A, 0-5s). Hunger. Hero opens an almost empty fridge. Discovers one lonely sausage. Despair flips to hope.
- Beat 2 (Clip B, 5-10s). Cooking. Pot, water, fire, sausage, bubbling, anticipation, hero shot lifting the sausage with a fork.
- Beat 3 (Clip C, 10-15s). Catastrophe. Sausage slips, falls, wet thud. Window. Bedroom. Hungry sleep. Cut to black.
Clip A. Hunger (production-grade skeleton applied)
Important direction:
This must be a clearly edited multi-shot sequence with visible cuts between shots.
Do not generate a single continuous take. Each shot must have a different camera angle
and different framing. Use rapid montage pacing.
Use @img1 as the main character reference and preserve the exact same man across the
whole clip: bald head, gray-green eyes, round expressive face, moustache, long black
braided goatee, short dark side hair tufts, slightly overweight build, tragicomic face.
Dress him in a dark oversized T-shirt and dark sweatpants. No nudity. No glasses.
No extra characters.
Generate a 5-second multi-shot tragicomic cinematic sequence with fast dynamic editing.
Story:
The man is hungry at night. He suddenly goes to the refrigerator, opens it, feels
disappointment because it is almost empty, then notices one lonely sausage and becomes
instantly hopeful.
Visual style:
Cold night kitchen. Blue-green refrigerator light. Desaturated colors. No warm yellow
tones. Slightly harsh LED reflections. Realistic cinematic look. High contrast but
natural skin texture. Subtle film grain. Realistic food detail.
Camera style:
Sony FX3 handheld look. 35mm and 50mm lenses for medium and close shots. 100mm macro
for inserts. Handheld micro-shake. Visible cuts between shots.
Editing style:
Fast montage with hard cuts. Each shot a different angle and framing. Tense rhythm,
escalating to the moment of discovery.
Audio:
Deep stomach growl, soft room tone, refrigerator hum, quiet footsteps, fridge door
sound. No dialogue. No subtitles. No on-screen text.
Shot 1, 0.0–0.8 sec. Extreme close-up of the man's eyes in darkness. He is awake,
hungry, tense. Static close shot, faint blue ambient light.
Shot 2, 0.8–1.6 sec. Medium handheld side shot. He sits up and walks fast to the
kitchen. Slight shake, push-in.
Shot 3, 1.6–2.6 sec. POV from inside the refrigerator. The door opens toward camera.
Cold blue-green light hits his face. Shelves almost empty. His face drops.
Shot 4, 2.6–3.5 sec. Rapid inserts: empty shelf, empty container, lonely sauce stain,
his sad eyes, his hand moving items aside.
Shot 5, 3.5–5.0 sec. Macro insert of one single sausage in the corner. Rack focus
from empty shelf to sausage. Smash cut to a slightly low-angle close-up of his face.
His eyes widen, despair flips to joy, tiny victorious smile.
Lighting:
Dark apartment, very low ambient fill. Main source is cold refrigerator light,
blue-green, top-front. Strong contrast. No warm kitchen light.
Composition:
Tight close-ups and inserts. Negative space inside the empty fridge. Final face shot
heroic and absurd.
--resolution 1080p --duration 5 --camerafixed falseClips B and C follow the same structure with their own story paragraph, shot list, and final image. The character lock and visual style blocks are repeated verbatim in each clip — Seedance has no memory between generations.
Older single-clip skeleton (kept for quick scenes)
For a single non-dramatic shot or a quick test, the original 6-step skeleton still works:
Subject. [identity block with face, hair, clothing, distinguishing features].
Motion. [one clear present-tense action for the scene].
Camera. Shot 1. [framing, lens, movement]. Cut to. Shot 2. [framing, lens, movement]. Cut to. Shot 3. [framing, lens, movement].
Environment. [location, time of day, props, weather].
Lighting. [dominant source, direction, quality, color].
Style. [realism level, genre reference, palette].
Audio. [ambient, SFX, silence moments]. (1.5+ only)
Continuity. [what must remain constant across shots].
--resolution 1080p --duration 5 --camerafixed falseFor dramatic, multi-shot, character-locked, or stitched-clip work — always use the 11-block production-grade skeleton from section 6.
Universal rules — apply to every video model
These rules apply to Seedance, Kling, Veo, and any other AI video generator. They exist because all current video models share a common failure mode. They reward concrete physical direction. They punish abstraction, contradictions, and keyword spam. Ignoring these produces muddy generations regardless of which model you target.
Read this file after dramaturgy.md and before any model-specific file. The model-specific syntax sits on top of these rules — it does not replace them.
Contents
1. The non-negotiable. Details intensify emotion (Details Law) 2. U1. Universal prompt skeleton 3. U2. Weight-at-start 4. U3. Show don't tell 5. U4. Natural language beats tag spam 6. U5. One primary camera move per shot 7. U6. Precise lens language 8. U7. Character consistency anchor 9. U8. No contradictions 10. U9. Concrete physical detail over abstract concept 11. U10. Duration discipline 12. U11. The final image rule 13. U12. The three-detail check (audit before sending)
---
1. The non-negotiable. Details intensify emotion. Laziness kills the prompt.
This is the most violated rule in AI video, and the one reason most multi-shot prompts fail. The dramaturgy is fine — the writer just got lazy on a single shot, and that one thin shot drags the whole sequence into mush.
Every shot in every prompt owns at least three concrete physical details:
1. One environmental pressure. Cold blue refrigerator light. Steam off boiling water. Wet asphalt. Flickering fluorescent. Dripping tap. Curtain breathing in the AC. (Kurosawa: weather is a character. See dramaturgy.md §9.) 2. One physical micro-action. Jaw locks. Knuckles whiten on the fork. Lips press into a line. He swallows hard. Fingers curl against the doorframe. (Show, not tell — the body is the only place where feelings render.) 3. One sound anchor or visual motif. Stomach growl at 2.3s. Reflection in a darkened phone screen. Rain on the same windowpane. A single fluorescent flicker before each cut.
If a shot has none of these — it is filler. Delete it or rewrite it. No exceptions for "establishing", "transition", or "hero product" shots — those are exactly the shots that go lazy first.
Words that do not render and mark the writer being lazy:
- "cinematic", "professional", "high quality", "masterpiece", "stunning", "epic", "amazing"
- "beautiful lighting", "dynamic camera", "intense moment", "powerful scene"
- "he is sad", "she is angry", "he is afraid" — emotions named without a body
Replace each with concrete physical facts. The full theory of why this works lives in dramaturgy.md §2 (Second law).
2. U1. Universal prompt skeleton
Build every prompt from these layers, roughly in this order:
[Subject / Character]
[Action / Motion]
[Scene / Environment]
[Camera / Shot / Lens]
[Lighting / Atmosphere]
[Style / Mood / Palette]
[Sound / Audio]
[Duration / Aspect ratio / Resolution]
[Continuity rules]
[Negative constraints — only if the model supports them]Model-specific skeletons live in their reference files (seedance.md, kling.md, veo.md). For dramatic / multi-shot / character-locked Seedance work, the production-grade 11-block skeleton in seedance.md §6 supersedes this generic skeleton.
3. U2. Weight-at-start
Generators put more attention on the first 30-40% of tokens. Lead with subject and action. Style modifiers go at the end. Camera, lighting, and environment live in the middle.
4. U3. Show don't tell
The model cannot render feelings. It renders bodies. Translate every emotion into a physical action.
- Bad. "He is scared."
- Good. "His jaw locks. He stops breathing for one beat. His fingers curl against the doorframe."
5. U4. Natural language beats tag spam
Video models are not image models. Tag stuffing like "masterpiece, 4k, cinematic, beautiful" fails. Write in full cinematic sentences as if briefing a human DOP.
6. U5. One primary camera move per shot
Do not stack three camera moves in a 5-second clip. Pick one dominant move (dolly-in, pan, tracking, static). Layer a subtle micro-adjustment if needed (slight handheld shake, gentle rack focus). More than that produces visual chaos.
7. U6. Precise lens language
State the lens. "Shot on 50mm" works across all major models. Quick map:
- 24mm. wide, immersive, exaggerated space
- 35mm. natural documentary
- 50mm. intimate, human perspective
- 85mm. portrait, compressed background
- 100mm macro. texture, detail
- anamorphic 40mm. cinematic widescreen
Full vocabulary in camera-lighting-vocabulary.md.
8. U7. Character consistency anchor
Anchor identity at the START of every prompt. In a multi-clip sequence, repeat the full identity block in every single prompt. Video generators have no memory between generations. Treat every clip like briefing a brilliant intern with memory damage.
Identity block must include: face shape, eye color, skin tone. Hair color, length, style. Facial hair. Exact clothing items. Distinctive accessories.
Model-specific syntax for character locking:
- Seedance:
@img1reference + identity block (seeseedance.md§10) - Kling 1.x – 2.x: Element Binding with 3-4 reference images (see
kling.md§7) - Kling 3.0: in-prompt
[Character A: <full identity>]labels, optionally combined with Element Library (seekling.md§3) - Veo: reference ingredients / JSON identity (see
veo.md)
9. U8. No contradictions
The model obeys the strongest signal. Contradictions produce artifacts.
- Bad. "Still pond" + "flowing water."
- Bad. "Close-up" + "wide cinematic landscape."
- Bad. "Quiet moment" + "explosive action."
10. U9. Concrete physical detail over abstract concept
"Loneliness" does not render. "A man sitting alone, shoulders collapsed, face lit by blue phone glow, empty bottles on the table" does.
This is the same rule as the Details Law in section 1, viewed from a different angle. Section 1 is the audit. This is the principle.
11. U10. Duration discipline
Most models work in 5-10 second clips. Longer narratives live in multiple clips stitched in the editor. Do not cram a 30-second story into a 5-second prompt.
Default splits:
- 10s = 2 clips × 5s
- 15s = 3 × 5s
- 30s = 6 × 5s
- 60s = 12 × 5s
Two exceptions:
- Seedance — 2-3 shots inside one 5-10s clip via "Cut to" syntax (see
seedance.md§8). - Kling 3.0 — up to 6 shots inside one generation, up to 15s, with native audio and dialogue (see
kling.md§3).
12. U11. The final image rule
Every clip needs a clear final frame. The model uses the ending as emotional destination.
- "Ends on his face frozen in the blue refrigerator light" beats "he stands there sadly."
The final image is also one of the five anchors (see dramaturgy.md §13). Naming it is non-negotiable.
13. U12. The three-detail check (audit before sending)
Before returning the final prompt to the user, audit every shot. Each shot must carry at least one of each:
1. Environmental pressure (lighting, weather, surface, sound of the room). 2. Physical micro-action on the body (jaw, hand, breath, eye, gesture). 3. Sound anchor or recurring visual motif tied to the emotional spine.
If a shot has zero, fix it before sending. If a shot has only one, ask whether you can make it two without bloat. The strongest prompts in this skill's worked examples always have all three.
Empty descriptors that fail this check: "establishing wide shot", "beautiful lighting", "dynamic camera move", "cinematic look", "intense moment", "dramatic close-up". Replace each with three concrete physical facts.
Veo reference (Google)
Contents
1. What Veo is 2. Versions and specs 3. Prompt structure and length 4. Dialogue syntax (critical, unique) 5. SFX syntax 6. JSON prompts (powerful, unique) 7. Image-to-video (Veo 3.1) 8. Reference ingredients (3.1) 9. Failure modes and fixes 10. Skeleton and example
---
1. What Veo is
Google's cinematic video model with NATIVE synchronized audio. The only major generator that creates dialogue, SFX, and music in sync with the video and does lip-sync natively. Best for commercial polish, narrative with dialogue, and cinematic audio.
2. Versions and specs
- Veo 3. Text-to-video, native audio, 4 / 6 / 8 second clips.
- Veo 3.1. Adds image-to-video with First Frame, improved audio, reference ingredients, stronger motion coherence.
Duration. 4, 6, or 8 seconds. Dialogue budget. Max 8 seconds of spoken audio per clip.
3. Prompt structure and length
Order matters. Lead with subject and camera. Quality modifiers go at the end.
[Subject / Action]
+ [Environment / Setting]
+ [Camera / Shot Type / Lens]
+ [Lighting / Atmosphere]
+ [Style / Quality]
+ [Audio]
+ [Duration]Sweet spot. 50-200 words.
- Shorter prompts = more creative latitude, less control.
- Longer prompts = tighter control, higher risk of contradictions.
4. Dialogue syntax (critical)
Veo is the only model that renders synced lip movement and voice. The syntax matters.
Required format
Double quotes + lead-in verb (says, whispers, shouts, mutters, asks).
A woman says, "Welcome to the future."
He whispers, "Don't move."
She shouts, "Get out!"Colon after the lead-in verb works too and is often more reliable.
A woman says: "Welcome to the future."Voice character modifiers
Place modifiers before the lead-in verb.
He says in a weary voice, "We are fine. We are fine."
She whispers nervously, "I don't want to be here."
He shouts excitedly, "We did it!"Timing rule
Maximum 8 seconds of spoken audio. Cram too many words into a 5s clip and the delivery speeds up unnaturally. Cut the line until it fits.
5. SFX syntax
Three supported formats. Mix as needed.
SFX: thunder cracks in the distance
Audio: rain on tin roof, distant traffic, one slow breath
(a loud thunderclap)
(key turning in a lock)
(wet footsteps on concrete)Use labels Audio:, Says:, SFX: to separate sound direction from visual direction. The model needs an explicit signal that audio should be generated.
6. JSON prompts (powerful, unique)
Veo parses structured JSON. This prevents "concept bleed" where describing mood accidentally changes object colors. Use JSON for complex scenes with strict continuity needs.
Full schema
{
"version": "veo-3.1",
"output": {
"duration_sec": 8,
"fps": 24,
"resolution": "1080p",
"aspect_ratio": "16:9"
},
"global_style": {
"look": "cinematic naturalism",
"color": "cold blue-gray palette, desaturated skin",
"mood": "quiet domestic tragedy",
"reference": "Fincher-style motivated camera"
},
"continuity": {
"characters": [
{
"id": "man",
"description": "40s, tired eyes, stubble, dark blue t-shirt, grey sweatpants, barefoot"
}
],
"props": ["empty refrigerator", "single sausage", "chipped white plate"],
"lighting_constant": "cold fridge light as key"
},
"scenes": [
{
"id": "01",
"start": "0.0",
"end": "3.0",
"shot": {
"type": "medium close-up",
"framing": "eye-level, slightly offset",
"camera": "slow push-in, 50mm"
},
"action": "He opens the fridge. His face catches the cold light. His eyes stop on the empty shelf.",
"environment": "small kitchen, 3am, rain outside window",
"lighting": "cold fridge light as key, warm window spill as rim",
"audio": "fridge hum, distant rain, one stomach growl"
},
{
"id": "02",
"start": "3.0",
"end": "8.0",
"shot": {
"type": "extreme close-up",
"framing": "macro insert on hand",
"camera": "static, 100mm macro"
},
"action": "His hand hovers over a single sausage. He picks it up slowly, exhales.",
"environment": "inside the fridge",
"lighting": "cold fridge light, high contrast",
"audio": "quiet breath, soft plastic crinkle"
}
]
}When to use JSON
- Multiple scenes in one generation.
- Strict character continuity across shots.
- Complex props that must stay the same color and size.
- When a prose prompt kept changing subject colors when you added mood words.
Not every Veo prompt needs JSON. For simple clips, prose is faster and often better.
7. Image-to-video (Veo 3.1)
Uses the static image as First Frame. Prompt guides motion and sound.
Rules.
- Do not re-describe static elements.
- Describe only motion, camera, light change, and audio.
- Add "maintain the subject from the first frame" to protect identity.
Example.
Maintain the subject from the first frame. Slow push-in, 50mm. She exhales, her eyes shift to the left, one strand of hair falls across her forehead. Warm rim light grows stronger.
Audio: soft breath, distant traffic, one door closing in the next room.
Duration: 6 seconds.8. Reference ingredients (3.1)
Upload multiple reference images (character, location, prop) and tag them in the prompt.
The character from reference_1 walks into the location from reference_2 holding the object from reference_3.Use when you need a specific character in a specific place with a specific object.
9. Failure modes and fixes
Dialogue speeds up unnaturally
Fix. Cut the line to fit 8 seconds of natural speech. Test by reading aloud.
Audio missing from output
Fix. Add explicit Audio:, SFX:, or Says: labels. Do not assume the model will infer audio from visual description.
Camera direction ignored
Fix. Lead the prompt with camera. "Wide aerial shot" at the start beats "cinematic camera work" in the middle.
Character color changes when mood words are added
Fix. Switch to JSON prompt. Lock the character description in the continuity block. Keep mood in global_style where it cannot bleed into object colors.
Lip-sync off
Fix. Check the lead-in verb. "She says, ..." outperforms just quoted speech alone. Colon form often more reliable than comma.
Prompt too long, model cherry-picks
Fix. Compress to the 50-100 word range. Or switch to JSON which handles length better because of structural separation.
10. Skeleton
Prose version
[Subject performing action] in [environment]. [Camera framing + lens + movement]. [Lighting direction + color temperature]. [Style + mood + palette].
Audio: [ambient sounds, SFX, music texture].
Says: [character] says, "[dialogue, max 8 seconds of speech]."
SFX: [punctual sound events].
Duration: [4 / 6 / 8] seconds.JSON version
See section 6. Copy the schema and fill in.
Worked example. Commercial hero shot
A woman in a cream silk blouse stands in front of a morning window, lifting a ceramic coffee cup to her lips. Medium close-up, 85mm, slow push-in. Warm window key from frame-left, soft bounce fill from frame-right. Cinematic naturalism, creamy palette, shallow depth of field.
Audio: distant city ambience, ceramic clink, one slow breath.
Says: She whispers to herself, "One more minute."
SFX: (spoon tapping ceramic at 2 seconds).
Duration: 6 seconds.