
Imagine Prompt
- 1 installs
- Updated May 18, 2026
- whitetowerai/imagine-skill
Rewrite image and video generation prompts for specific Vofy models using model-specific official prompt guides.
About
Optimizes user intent into model-ready prompts for Vofy image and video models, loading per-model guides and preserving creative direction. A developer uses it to improve, rewrite, translate, or adapt a media prompt before generating.
- Model guide routing table for sora, gpt-image, gemini, and veo models
- Output format gives optimized prompt plus parameter hints and next command
Imagine Prompt by the numbers
- 1 all-time installs (skills.sh)
- Ranked #1,200 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Jul 23, 2026 (Skillselion catalog sync)
npx skills add https://github.com/whitetowerai/imagine-skill --skill imagine-promptAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| Last updated | May 18, 2026 |
| Repository | whitetowerai/imagine-skill ↗ |
What it does
Rewrite image and video generation prompts for specific Vofy models using model-specific official prompt guides.
Files
Prompt Optimization For Vofy Media
Rewrite user intent into model-ready prompts while preserving creative direction.
Adaptive Workflow
1. Identify output type, mode, named model, source media, target platform, aspect ratio, duration, and hard constraints. 2. If no model is named, choose or infer one with imagine-models before loading prompt guidance. 3. Load the exact model guide from the routing table below; for multi-stage pipelines, load one guide per stage. 4. If the selected model has no guide in this skill, do not load a substitute guide; use only the user's intent plus imagine-models capability constraints. 5. Rewrite the prompt for the selected model and mode. 6. When the user asked to generate media, pass the optimized prompt to imagine-create instead of stopping at advice.
Model Guide Routing
| Model | Load |
|---|---|
sora-2, sora-2-pro | references/sora-2.md |
gpt-image-1.5 | references/gpt-image-1.5.md |
gpt-image-2 | references/gpt-image-models.md |
gemini-2.5-flash-image, gemini-3-pro-image-preview, gemini-3.1-flash-image-preview | references/gemini-image.md |
veo-3.1, veo-3.1-fast, veo-3.1-lite | references/veo.md |
Rewrite Shape
For images, include: subject, environment, composition, style or medium, lighting, camera or render traits, color palette, important text, and constraints.
For videos, include: subject, action over time, setting, camera movement, shot scale, pacing, lighting, style, duration-aware beats, continuity constraints, and audio/dialogue only if the chosen model supports it.
For edits or source-driven modes, describe what must remain unchanged before describing the change.
Output Format
- Optimized prompt: one copy-paste-ready prompt.
- Parameter hints: model, mode, ratio, duration, resolution, and special flags when known.
- Why this works: 1-3 terse bullets only when useful.
- Next command: include a
vofycommand only if the user asked to create media.
Guardrails
- Do not invent unsupported flags; verify special controls with
vofy models <model>orimagine-modelsreferences. - Do not use another model's prompt guide as a fallback.
- Keep user-specified text exact, especially logos, UI copy, captions, product names, and brand language.
- Do not overconstrain simple prompts; add detail where it improves controllability.
Gemini Image Prompt Guide
Source: https://ai.google.dev/gemini-api/docs/image-generation?hl=zh-cn#prompt-guide
Applies to: gemini-2.5-flash-image, gemini-3-pro-image-preview, gemini-3.1-flash-image-preview.
Primary Prompting Principle
- Write prompts as complete, descriptive instructions; Gemini responds better to clear natural language than to comma-separated tags.
- Include intent and context, not just visual attributes: what the image is for, who or what matters, and what should be easy to recognize.
- Be specific about the visual outcome: subject, action, environment, composition, lighting, mood, palette, material, texture, and text.
- Use ordered steps for complex edits, multi-image composition, diagrams, or sequential scenes.
- Prefer positive descriptions of the desired scene over long negative lists; use explicit exclusions only for critical failure cases.
- Put model capabilities in Vofy flags when available: model, mode, input images, aspect ratio, resolution, reasoning, search, and output handling.
Prompt Anatomy
Create [image type] for [purpose/context].
Show [specific subject], [action/expression], set in [environment/background].
Use [composition/shot/camera angle], [lens or render traits], [lighting], [mood], and [palette].
Emphasize [materials, textures, details, product/brand constraints].
Text: [exact quoted text, font description, placement, or "no text"].
Preserve/change: [reference-image or edit constraints].
Output: [layout, aspect ratio, or platform need when not handled by flags].
Avoid: [only critical exclusions].Generation Prompt Patterns
Photorealistic People Or Scenes
Use when the target should look like a real photo.
Include:
- Subject identity, age range or role, expression, pose, and action.
- Location, time of day, weather, props, and background details.
- Shot type, camera angle, focal length or depth of field, and framing.
- Lighting type, light direction, mood, palette, and level of realism.
- Texture details such as skin, fabric, metal, glass, dust, water, or plants.
Template:
Create a photorealistic [shot type] of [subject] [action/expression] in [setting]. Use [camera angle/lens/framing], [lighting], [mood], and [palette]. Emphasize [specific textures/details]. No text.Stylized Illustration
Use when the target should be clearly drawn, painted, rendered, or designed.
Include:
- Medium: watercolor, vector, ink, clay render, 3D toy, paper cutout, pixel art, isometric, editorial illustration.
- Shape language: rounded, geometric, delicate, chunky, simplified, ornate.
- Line quality, brush texture, shading, and detail level.
- Palette, background, composition, and intended use.
Template:
Create a [medium/style] illustration of [subject] for [use]. Use [shape language], [line/brush/render traits], [palette], and [composition]. Keep the background [simple/specific]. No text unless specified.Stickers, Icons, Mascots, And Assets
Use for isolated assets that may be cut out or placed into apps.
Include:
- A single strong subject with readable silhouette.
- Asset style, outline, shadow, palette, and background.
- Facial expression or brand personality for mascots.
- White or plain background; do not rely on transparent-background wording unless the model/tool explicitly supports it.
Template:
Create a [sticker/icon/mascot] of [subject] with [expression/personality]. Use [style], clean outlines, simple readable shapes, [palette], and a plain white background. Center the subject with generous padding. No text.Text, Logos, Posters, And Layouts
Use when text accuracy or visual hierarchy matters.
Include:
- Exact text in quotation marks.
- Font description rather than obscure font names.
- Placement, hierarchy, alignment, spacing, and contrast.
- Logo mark concept, brand mood, palette, and background.
- “No extra text” when accuracy matters.
Template:
Create a [logo/poster/layout] for [brand/concept]. Render exactly the text "[TEXT]" in [font style], placed [position]. Use [visual style], [palette], [layout constraints], and high contrast. No extra text.Product And Commercial Photography
Use for ads, catalog shots, packaging, product pages, and mockups.
Include:
- Product shape, material, color, finish, label/logo preservation, and visible features.
- Surface, background, props, and scale cues.
- Lighting setup, reflection, shadow, camera angle, and focus.
- Commercial mood: premium, playful, clinical, editorial, lifestyle, luxury, tech.
Template:
Create a professional product photograph of [product] for [use]. Show [geometry/material/features] on [surface/background]. Use [camera angle], [lighting setup], crisp focus on [feature], realistic shadows/reflections, and a [brand mood] palette. Preserve [logo/text/details].Minimal Designs And Negative Space
Use when future copy or UI elements need space.
Include:
- Exact subject placement and size in frame.
- Empty region location and style.
- Background color, gradient, texture, or environment.
- Mood, lighting, and platform use.
Template:
Create a [style] image for [platform/use]. Place [subject] in [frame position/size]. Leave a large clean empty area on [side/top/bottom] for future text. Use [background], [lighting], and [palette]. No text.Diagrams, Infographics, And Educational Images
Use when the image must explain a concept clearly.
Include:
- Diagram style and audience level.
- Limited, exact labels.
- Spatial relationship, arrows, callouts, color coding, and hierarchy.
- Readability, high contrast, and clean background.
Template:
Create a clean [diagram/infographic] explaining [concept] for [audience]. Show [components] arranged [layout]. Add exactly [number] labels: "[A]", "[B]", "[C]". Use [arrows/callouts/color coding], high contrast, and a plain background. No extra text.Sequential Art And Storyboards
Use for comics, panel sequences, process visuals, and shot boards.
Include:
- Number of panels and reading direction.
- Consistent character description and style.
- Per-panel action beats and any captions.
- Continuity constraints for clothing, props, location, and lighting.
Template:
Create a [number]-panel [comic/storyboard] in [style]. Keep [character] consistent across all panels: [appearance/clothing]. Panel 1: [beat]. Panel 2: [beat]. Panel 3: [beat]. Use [layout], [palette], and [caption/text rules].Editing Prompt Patterns
Add, Remove, Or Modify A Specific Element
Use for targeted changes to an existing image.
Include:
- What the source image contains.
- The exact change.
- What must stay identical.
- Matching style, lighting, shadows, perspective, texture, and scale.
Template:
Using the provided image of [source], change only [target element] to [new description]. Keep [unchanged elements] exactly the same. Match the existing [style/lighting/shadows/perspective/texture].Local Repainting Or Masked Edit
Use when a mask or a clearly localized region exists.
Include:
- The affected region in words even if a mask is present.
- The desired replacement.
- Strong “keep everything else unchanged” instruction.
- Integration requirements for edges, shadows, reflections, and perspective.
Template:
Edit only the [masked/specified region]: [change]. Preserve every unmasked part of the image, including [identity/composition/background/lighting/camera]. Blend the edit naturally with matching shadows, edges, and perspective.Style Transfer
Use to re-render an existing scene in a new visual style.
Include:
- Preserve composition, subject layout, pose, and key identity details.
- Target medium, rendering style, brushwork, palette, lighting, and texture.
- What not to reinterpret: faces, products, logos, text, or proportions.
Template:
Transform the provided image into [target style]. Preserve the original composition, subject placement, pose, and key details. Render it with [medium/brush/render traits], [palette], and [lighting]. Do not alter [identity/logo/text/proportions].Sketch, Wireframe, Or Rough Concept Refinement
Use to turn drafts into polished visuals.
Include:
- What the sketch represents.
- Which sketch features must remain.
- Desired final material, style, detail level, lighting, and environment.
- Any labels or UI details that must remain exact.
Template:
Refine this rough [sketch/wireframe] into a polished [final style]. Preserve [layout/shape/key markings]. Add [materials/details/environment/lighting] while keeping the original concept recognizable. Keep any visible text exactly as provided.Character Consistency
Use for new poses, views, scenes, or expressions using an existing character image.
Include:
- Reference image role for character identity.
- Distinctive traits to preserve.
- New pose, camera angle, expression, clothing change, or environment.
- If pose is difficult, provide or mention a separate pose reference and assign its role.
Template:
Use image 1 as the character identity reference. Preserve [face shape, hairstyle, outfit, colors, proportions, distinctive marks]. Create a new image of the same character [new pose/action/view] in [setting]. Match the original design style while changing only [intended changes].High-Fidelity Preservation
Use when exact likeness, product marks, labels, UI, or geometry are critical.
Include:
- The high-value details that must remain exact.
- The specific allowed change.
- Instruction to avoid reinterpretation or simplification of those details.
Template:
Preserve [face/logo/product shape/label/UI text] with high fidelity. Do not redraw, simplify, or reinterpret those details. Change only [allowed target] to [new result], matching the original image’s camera, lighting, perspective, and texture.Multi-Image Composition
Use when multiple references provide identity, style, layout, background, pose, or product details.
Template:
Use image 1 as [identity/product/layout source].
Use image 2 as [style/background/material/pose reference].
Use image 3 as [optional additional role].
Create [final image description].
Preserve [critical identity/logo/text/proportions/composition].
Make the result unified by matching [lighting, shadows, perspective, color temperature, texture, scale].Text Rendering Guidance
- Quote exact text and keep it short.
- Describe the font style in plain visual terms: clean bold sans-serif, elegant serif, handwritten marker, rounded playful letters, condensed editorial type.
- Specify placement, alignment, size hierarchy, contrast, and surrounding whitespace.
- For logos, define the relationship between wordmark and symbol.
- For diagrams, use a small number of labels and explicit arrows or callouts.
- Add “no extra text” when text accuracy matters.
- For best results on text-heavy visuals, write the text content first, then ask for a visual layout that renders that exact copy.
- Prefer
gemini-3-pro-image-previewfor professional assets, complex text, or high-fidelity typography when available.
Semantic Negative Prompting
Gemini generally responds better when the target state is described positively.
Prefer:
Create an empty modern kitchen countertop with a clean marble surface, no objects on the counter, and soft morning light.Instead of only:
No cups, no plates, no appliances, no clutter.Use direct exclusions when they are essential, but pair them with a clear positive scene description.
Model-Specific Capability Notes
gemini-2.5-flash-image: fast interactive image generation and editing; useful for rapid iteration, conversational edits, and small reference sets.gemini-3-pro-image-preview: stronger for professional visuals, richer reasoning, high-fidelity text, up to 4K output, and larger multi-image reference workflows.gemini-3.1-flash-image-preview: focused image generation/editing preview with strong likeness/detail preservation in supported workflows.- Image generation accepts text and images as inputs; audio and video inputs are not supported for these image workflows.
- Requested image counts may not be returned exactly; design prompts so one strong output is acceptable unless the CLI/model explicitly supports multiple outputs.
- Generated images include SynthID watermarking.
Search And Reasoning Hints
- Use search-capable flags for current products, places, landmarks, wildlife, events, weather, score graphics, or factual visual references.
- Do not rely on Google Search grounding for real-world images of people with
gemini-3.1-flash-image-preview. - Use reasoning-capable flags for complex layouts, multi-reference composition, diagrams, precise product constraints, and high-fidelity text.
- Keep search/reasoning as execution parameters; the prompt should still specify the desired final visual result.
Input Reference Limits
Check imagine-models or vofy models <model> for the current source of truth. As of this guide:
gemini-2.5-flash-image: up to 3 input images.gemini-3-pro-image-preview: up to 5 high-fidelity input images and up to 14 total images.gemini-3.1-flash-image-preview: can preserve likeness for up to 4 characters and detail fidelity for up to 10 objects in one workflow.
Vofy Parameter Hints
- Put visual intent in
--prompt. - Use
--imagefor source images or references when the selected mode supports it. - Use aspect ratio and resolution flags rather than burying hard technical output constraints only in prose.
- Use search/reasoning flags only if the selected Vofy model exposes them.
- Do not invent flags; verify with
vofy models <model>orskills/imagine-models/.
Examples
Photorealistic Portrait
Create a photorealistic close-up portrait for an artisan profile. Show an elderly Japanese ceramicist with deep sun-etched wrinkles and a warm, knowing smile as he inspects a freshly glazed tea bowl in a rustic, sunlit workshop. Use a vertical portrait composition, soft golden-hour window light, an 85mm portrait lens with gentle bokeh, and a serene masterful mood. Emphasize clay texture, apron fabric, shelves of pottery, and warm earth tones. No text.Stylized Illustration
Create a whimsical watercolor illustration for a children’s book cover. Show a small fox wearing a navy scarf, standing under giant glowing mushrooms in a rainy forest. Use soft brush edges, gentle ink outlines, a teal and amber palette, visible paper texture, and a magical cozy mood. Leave clean space at the top for a future title. No text.Sticker Asset
Create a cute sticker of a sleepy orange tabby cat curled around a tiny laptop. Use rounded shapes, thick white outline, soft cel shading, warm pastel colors, and a plain white background. Center the sticker with generous padding. No text.Product Photography
Create a high-resolution studio product photograph for a minimalist coffee brand. Show a matte ivory ceramic mug on a pale stone surface, with the handle turned 45 degrees toward camera. Use a three-point softbox lighting setup, a low three-quarter camera angle, crisp focus on the ceramic rim, soft contact shadow, and a subtle reflection. Clean warm-gray background, premium catalog style. No text.Logo With Text
Create a modern minimalist logo for a coffee shop called "The Daily Grind". Render exactly the text "The Daily Grind" in a clean, bold sans-serif style inside a simple circle. Use black and white only, integrate a coffee bean shape in a clever but readable way, centered composition, high contrast. No extra text.Diagram With Labels
Create a clean flat-vector diagram explaining how a plant absorbs water for elementary students. Show soil, roots, stem, and leaves arranged vertically. Add exactly three large labels: "Roots", "Stem", "Leaves". Use blue arrows moving from soil through roots to leaves, high contrast, and a white background. No extra text.Local Edit
Using the provided image of the living room, change only the blue sofa to a vintage brown leather chesterfield sofa. Keep the pillows, wall art, floor, windows, camera angle, lighting, shadows, and overall composition exactly the same. Match the new sofa to the room’s perspective and soft daylight.Style Transfer
Transform the provided city street photo into a detailed ink-and-watercolor travel illustration. Preserve the street layout, buildings, pedestrians, and camera angle. Use delicate black linework, loose transparent washes, warm afternoon light, and subtle paper texture. Do not change storefront signs or visible text.Multi-Image Product Composite
Use image 1 for the exact perfume bottle shape, cap, label, logo placement, and product proportions. Use image 2 only for the warm golden studio lighting and satin background texture. Create a premium centered product ad with a soft reflection beneath the bottle, realistic shadows, cream-gold palette, and sharp label detail. Preserve all label text; no extra text.Character Consistency
Use image 1 as the character identity reference. Preserve the character’s round face, short silver hair, teal jacket, yellow scarf, and compact proportions. Create a new three-quarter rear view of the same character looking over a snowy mountain valley at sunrise. Keep the same stylized 3D toy render style and soft warm lighting.Sequential Panels
Create a 4-panel silent comic in a clean pastel vector style. Keep the same small robot character in every panel: round head, blue body, single antenna, expressive eyes. Panel 1: the robot finds a wilted plant. Panel 2: it brings a tiny watering can. Panel 3: the plant grows a flower. Panel 4: the robot smiles beside the flower. Use consistent lighting and no text.Semantic Negative Prompt
Create a cinematic wide-angle photo of an empty desert highway at dawn. The road stretches into open sand with no signs of traffic, no parked vehicles, and no buildings on the horizon. Use low warm sunlight, long shadows, dusty atmosphere, and a quiet isolated mood.Common Failure Fixes
- Prompt is too tag-like: rewrite it as a descriptive paragraph with subject, action, setting, lighting, camera, mood, and purpose.
- Result ignores reference images: assign every image a role and state what each one controls.
- Text is wrong: shorten copy, quote exact text, describe font and placement, and request no extra text.
- Layout is messy: reduce object count, specify positions, and use step-by-step instructions.
- Edit drifts too much: begin with what must remain unchanged, then describe the single allowed change.
- Unwanted objects appear: describe the desired clean/empty state and add only the most important exclusions.
- Style is inconsistent: specify medium, line quality, rendering method, palette, and detail level.
- Product or face changes: explicitly preserve identity, geometry, logos, labels, proportions, camera, and lighting.
Final Rewrite Checklist
- The prompt reads like a clear instruction or scene description, not a tag list.
- Purpose/context is explicit when it affects design decisions.
- Subject, action, setting, composition, lighting, mood, textures, and palette are concrete.
- Reference-image roles and preservation rules are named when inputs exist.
- Exact text is quoted, minimal, and has placement/font guidance.
- Complex scenes or edits are broken into ordered steps.
- Negative constraints are framed as desired visual states where possible.
- Search/reasoning/ratio/resolution controls are handled through Vofy flags when supported.
GPT Image 1.5 Prompt Guide
Source: https://developers.openai.com/cookbook/examples/multimodal/image-gen-1.5-prompting_guide
Applies to: gpt-image-1.5.
Model Behavior
- Use direct natural language instructions instead of keyword piles.
- State the task type first: generate, edit, inpaint, product mockup, text rendering, or style transfer.
- Be explicit about preservation during edits; the model may otherwise reinterpret the whole image.
- Put exact rendered text in quotes and keep it short.
- Separate visual description from API or CLI parameters; put size, background, number of images, image input, mask input, and download behavior in Vofy flags.
- Prefer the model's default behavior unless the user needs a specific aspect ratio, transparent background, exact text, or localized edit.
- When quality matters, ask for observable image traits rather than abstract intent: material, geometry, camera angle, lighting, shadows, palette, typography, and whitespace.
Core Prompt Anatomy
Goal: [generate/edit/inpaint] [asset type] for [use case].
Subject: [main subject with concrete details].
Composition: [viewpoint, crop, layout, whitespace, background].
Style: [medium, realism level, design style, era, texture].
Lighting/color: [light source, mood, palette].
Text: [exact quoted text, typography, placement, or "no text"].
Constraints: preserve [attributes]; avoid [critical exclusions].For simple generation, combine these into one concise paragraph. For edits, keep sections separate so preservation and changes are unambiguous.
Use this structure as a checklist, not a template to expose to end users. Drop sections that do not affect the result.
Generation Guidance
- Lead with the deliverable: poster, icon, product photo, UI illustration, diagram, character sheet, sticker, texture, logo draft.
- Specify visible layout rather than intent: “centered product with 30% empty space on the right,” not “premium.”
- Define style with medium and texture: studio product photo, flat vector illustration, risograph poster, clay render, watercolor sketch.
- Mention background and transparency needs explicitly, and use
--background transparentwhen supported. - For text-heavy outputs, reduce copy and use large, simple typography.
Composition And Camera
- State the camera relationship: eye-level, overhead flat lay, macro close-up, three-quarter front view, isometric, wide establishing shot.
- Define framing and crop: full body, waist-up, centered product, edge-to-edge pattern, safe margins, negative space.
- Give spatial instructions for multi-object scenes: left/right placement, foreground/background, overlap, scale hierarchy, and clear separation.
- For product or UI work, specify alignment, surface, reflection, shadow, and whether the asset should feel photographed, rendered, or illustrated.
- Use aspect ratio and resolution flags for canvas shape; use prompt text for what occupies the canvas.
Style Control
- Combine medium, era, finish, and constraints:
flat vector mascot, 1960s travel-poster palette, subtle paper grain, no gradients. - Avoid vague style-only prompts such as
make it cinematic; pair mood words with concrete lighting, lens, palette, and composition. - If matching a style reference, name the transferable traits: line weight, color palette, texture, lighting, typography, or composition.
- If style must not drift, add exclusions for competing looks:
not 3D, not photorealistic, no glossy plastic.
Editing Guidance
Use this order:
Preserve: [identity, pose, geometry, lighting, camera angle, colors, text/logo].
Change: [one primary edit].
Blend: match [shadows, reflections, grain, perspective, edge softness].
Avoid: [unwanted changes].- For masks, describe the masked region and the replacement content.
- For localized edits, use one primary change per request and say
onlywhen scope matters. - For product or logo work, preserve geometry, label text, logo placement, and brand colors.
- For people, preserve identity, expression, age, hairstyle, and clothing unless those are the edit target.
- If an edit fails, reduce the requested change to one local operation.
Inpainting With Masks
- Treat the mask as the edit boundary, but still describe what belongs inside it.
- Explain how the new content should connect to the unmasked image: perspective, edge softness, material, shadows, reflections, lighting direction, and grain.
- Preserve unmasked regions explicitly when important:
Keep all unmasked pixels visually unchanged. - If replacing a background, describe both the new background and how foreground edges should blend.
Prompting For Text
- Quote exact text:
Text: "SPRING SALE". - Keep text short and high contrast.
- Specify font style and placement: bold condensed sans-serif, top-left, centered baseline, large readable letters.
- Avoid asking for paragraphs, tiny captions, dense menus, or many labels in one image.
- Add
No extra text, no watermark, no misspellingswhen text accuracy matters. - For logos, signage, packaging, and UI, identify which text must remain exact and which decorative text should be omitted.
Reference Images
- Use references to anchor identity, product details, composition, or style.
- Say which attributes to copy and which to ignore; references do not replace written instructions.
- Do not assume the model knows which reference is primary; name the role of each input if multiple images are supported.
- When using multiple references, assign a stable role to each image before giving the edit or generation instruction.
- Use references for one or two strong anchors, not as a substitute for resolving contradictory requirements.
Use image 1 for the product shape and label placement. Use image 2 only for the warm studio lighting style. Preserve the exact bottle silhouette and front logo; replace the background with matte cream paper.Parameter Hints
- Use
--sizeor ratio-related Vofy options for canvas dimensions instead of asking for pixel dimensions in the prompt. - Use
--background transparentfor transparent assets and still mention transparent background in the prompt. - Use
--imageinputs for references or edits, and--maskonly when the user needs an inpainted region. - Use
--nfor variations rather than asking the model to create contact sheets or multiple unrelated designs in one image. - If Vofy exposes a quality flag for the selected model, use lower quality for draft composition checks and higher quality for final text, product detail, or subtle edits.
- Use
--yesin create examples so agents can run non-interactively.
Use Case Recipes
- Infographics and diagrams: define the audience, topic, major sections, flow direction, label count, visual hierarchy, and exact labels. Ask for simple readable structure, not dense copy.
- Image translation/localization: preserve layout, icons, colors, typography style, spacing, and imagery. Change only the specified text into the target language.
- Natural photorealism: write like a photographer: real subject behavior, camera/lens feel, shot scale, lighting source, depth of field, material wear, skin or fabric texture, and no heavy retouching.
- World-knowledge scenes: include the place, date or era, cultural context, and realism target. Ask for period-accurate clothing, objects, environment, and signage only when needed.
- Logos and marks: request an original non-infringing mark with simple geometry, strong silhouette, balanced negative space, flat color, scalability, centered padding, and no watermark.
- Comics and panels: specify panel count, reading order, equal panel sizing, one concrete visual beat per panel, repeated character traits, and any text or caption limits.
- UI mockups: describe a real shipped interface: platform, screen type, layout hierarchy, sections, spacing, typography, component states, practical content, and minimal decoration.
- Style transfer: preserve only style traits such as palette, texture, brushwork, line weight, lighting, or grain. Replace the subject or scene explicitly.
- Virtual try-on: lock identity, face, skin tone, body shape, pose, hair, expression, camera angle, and background. Change only garments and require realistic fit, folds, seams, occlusion, and shadows.
- Sketch-to-render: preserve the sketch layout, proportions, perspective, and design intent. Add realistic materials, lighting, surfaces, edges, and environment details.
- Product mockups: preserve product geometry, label text, logo placement, proportions, and material. For cutouts, request crisp alpha edges, no halos, and optional subtle shadow.
- Marketing creatives: separate product preservation from scene generation. Put final ad copy in quotes, specify exact placement, typography, contrast, and forbid extra characters.
- Lighting/weather changes: change only time of day, season, sky, weather, reflections, and color temperature. Preserve scene geometry, camera, subject placement, and identity.
- Object removal or recolor: name the object and operation, say
Do not change anything else, and require background reconstruction, matching texture, shadows, and lighting. - Person insertion or compositing: name the source person and target scene roles. Match scale, perspective, contact shadows, lighting direction, grain, and occlusion.
- Multi-image compositing: assign each image a role, then describe the final composition, which traits to preserve, which to ignore, and how sources should interact.
- Interior design swaps: preserve room architecture, camera, windows, floor, wall boundaries, and lighting. Change only specified furniture, materials, colors, or decor.
- Merch concepts: describe product type, material, pose, package or accessory details, scale cues, clean product render lighting, and plain background.
- Character-consistent book art: create or reuse a character sheet, then lock character traits across scenes while varying pose, emotion, setting, and story beat.
Useful Patterns
Product Photo
Generate a studio product photo of a matte white insulated tumbler, three-quarter front view, centered on a light gray seamless background. Softbox reflection from upper left, subtle contact shadow, crisp rim highlight, neutral premium catalog style. No text.Product Background Edit
Preserve the product silhouette, label text, logo placement, camera angle, and existing highlights. Replace only the background with a warm beige stone countertop and soft morning window light from the left. Match the contact shadow, reflections, perspective, and edge softness. Do not change the product color, cap, label, or text.Transparent Sticker
Generate a cute flat-vector sticker of a smiling corgi astronaut floating with a tiny blue planet. Thick white sticker border, clean shapes, cheerful orange and sky-blue palette, transparent background. No text.Local Edit
Preserve the person, pose, camera angle, and warm indoor lighting. Replace only the plain gray sweater with a forest-green cable-knit sweater. Match fabric folds and shadows. Do not change the face, hair, hands, or background.Masked Inpaint
Edit only the masked area on the tabletop. Add a small ceramic espresso cup in correct perspective, with a soft contact shadow and matching warm indoor light. Keep all unmasked regions visually unchanged, including the laptop, notebook, hands, and background.Text Poster
Generate a minimalist concert poster on an off-white paper background. Large black text at top: "MOON ROOM". Smaller text below: "LIVE 9 PM". Center: simple blue crescent moon icon, lots of whitespace, modern Swiss grid layout. No extra text.Style Reference
Use the reference image only for its muted teal-and-cream palette, screen-print texture, and simple geometric shapes. Generate a new poster of a mountain train crossing a bridge at sunrise, centered composition with large empty sky at the top. Do not copy the reference subject, layout, logo, or any text. No text.Common Failure Fixes
- Wrong edit scope: start with
Preserve:and list what must not change. - Weak composition: specify crop, camera angle, object placement, and whitespace.
- Messy text: shorten text, quote it, increase size, and remove extra labels.
- Style drift: name medium, era, texture, palette, and lighting.
- Overloaded request: split into separate generations or edits.
- Reference confusion: assign each reference a role and state which traits to ignore.
- Poor background removal: request transparent background in both prompt and Vofy flags.
Final Rewrite Checklist
- The prompt states generation, editing, or reference use.
- The main subject and deliverable are explicit.
- Composition, style, lighting, and background are concrete.
- Preservation instructions come before edit instructions.
- Exact text is quoted, short, and placed.
- Vofy flags carry size, background, image, and mask settings.
- Reference-image roles and ignored traits are explicit when references are used.
- Masked edits describe blending with surrounding perspective, light, shadows, and texture.
GPT Image Models Prompt Guide
Source: https://developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide
Applies to: gpt-image-2 and general OpenAI image-generation prompting.
Core Rules
- Write prompts as clear instructions, not comma-separated keyword lists.
- Put the final asset type first: hero image, product mockup, icon, diagram, edit, inpaint, style transfer, poster, storyboard frame, or transparent cutout.
- Describe the visible result, not hidden intent: subject, setting, composition, style, lighting, color, text, and constraints.
- Use prompt text for creative direction; use Vofy flags for model, size, quality, background, images, masks, counts, and output handling.
- For source-image work, separate what to preserve from what to change before adding style or polish.
- Use explicit tradeoffs: exact instruction following, polished aesthetics, speed, cost, or reference fidelity.
Model And Parameter Choices
- Use
gpt-image-2when the prompt needs high instruction adherence, precise edits, reliable text, world knowledge, or strict image-input fidelity. - Use
gpt-image-1.5for fast drafts, simple edits, and lower-cost iteration before final rendering. - Use low quality for composition exploration, medium for normal iteration, and high for final outputs with product detail, readable text, or subtle edits.
- Pair transparent-background output with prompt constraints: isolated subject, clean alpha edges, no backdrop, controlled shadows, and no unwanted halo.
- Move operational settings out of the prompt: size, quality, transparency, masks, and image references belong in Vofy parameters when available.
Prompt Anatomy
Create [asset type] for [use case/audience/platform].
Subject: [main subject, identity, attributes, materials].
Scene: [environment, background, props, context].
Composition: [crop, viewpoint, placement, negative space, aspect-ratio-aware framing].
Style: [medium, realism level, genre, era, texture, render treatment].
Lighting/color: [light source, mood, palette, contrast].
Text: [exact quoted copy, hierarchy, placement, typography, or "no text"].
Constraints: [preserve, avoid, safety margins, brand rules, output-specific requirements].Keep simple tasks short. Add structure only when precision matters.
Task Patterns
New Image
- State the deliverable and audience first.
- Anchor the main subject with concrete visible attributes.
- Add composition, camera, background, lighting, style, and color in that order.
- End with text and avoid-list constraints.
Image Edit
Preserve: [identity, shape, pose, camera angle, perspective, lighting, shadows, logo/text, background details].
Edit: [one primary local change].
Integrate: match [materials, reflections, grain, edge softness, color temperature, depth of field].
Avoid: [unwanted changes].- Put preservation before the edit request.
- Avoid combining unrelated edits; inspect after one precise change, then iterate.
- For masks, describe the masked region and replacement content in visual terms.
- For identity-sensitive edits, repeat the attributes that must not change.
Reference Remix
- Assign a role to every input: identity, product geometry, pose, layout, palette, material, lighting, or style.
- Name the primary reference when multiple images are supplied.
- State what to ignore from each reference, especially unwanted background, text, logos, crop, or pose.
- If style transfer should preserve content, say “style only; do not change subject, pose, layout, or text.”
- If content should change but style should remain, isolate style traits such as palette, lens, brushwork, paper texture, or lighting.
Text Rendering
- Quote exact words and keep them short.
- Specify hierarchy: title, subtitle, label, small caption, badge, or UI copy.
- Specify typography visually: bold geometric sans-serif, serif editorial headline, hand-lettered marker, monospaced UI label.
- Specify placement and contrast: top-left, centered, lower third, large readable letters, high contrast against plain background.
- Ask for no extra text, no watermark, and no misspellings when text accuracy matters.
- For dense information, first generate a layout with placeholder blocks, then iterate on final copy.
Product And Brand Assets
- Preserve product geometry, silhouette, label layout, logo placement, visible text, and brand colors.
- Specify material behavior: matte plastic, brushed metal, glass thickness, leather grain, paper texture, soft fabric weave.
- Lock camera angle and lighting for mockups: three-quarter front view, top-down flat lay, macro close-up, softbox reflection.
- Ask for realistic contact shadows, reflections, scale, and edge sharpness.
- Avoid redesigning labels or changing brand colors unless the user asks.
Transparent Assets
- Describe silhouette, contour quality, internal highlights, and shadow behavior.
- Use: isolated subject, clean alpha edges, no backdrop, no halo, no rectangular canvas, no ground shadow unless requested.
- For stickers, specify border thickness, cut line, flat/vector style, and transparent outside border.
- For icons, specify centered composition, simple shapes, readable at small sizes, and no tiny details.
Diagrams And Infographics
- Prioritize readability over decoration.
- Specify clean layout, large labels, arrows, hierarchy, and color coding.
- Limit the number of labels; prefer fewer large labels over many tiny captions.
- Use exact terms in quotes and request no extra text.
- For step diagrams, state order and flow direction: clockwise, left-to-right, top-to-bottom, or numbered stages.
UI, App, And Screen Mockups
- Specify device, screen type, layout density, spacing, and visual system.
- Quote UI labels exactly and keep them few.
- Ask for consistent alignment, readable text, realistic shadows, and no invented extra UI copy.
- For reference-driven UI, preserve component placement, brand colors, icon shapes, and hierarchy.
Characters And Portraits
- Specify age range, expression, pose, clothing, hairstyle, gaze direction, camera angle, and lighting.
- For edits, preserve identity, face shape, expression, skin tone, hairstyle, pose, and clothing unless those are the target.
- Avoid ambiguous identity changes by making the edit local and concrete.
- Use lens and framing controls: headshot, waist-up, full-body, eye-level, 85mm portrait, shallow depth of field.
Realistic Photography
- Include camera viewpoint, focal length feel, depth of field, lighting source, lens artifacts only if useful, and environment detail.
- Use natural constraints: physically plausible shadows, reflections, contact points, scale, and perspective.
- Avoid overloading with conflicting styles such as “photorealistic watercolor vector.”
Stylized Illustration
- Name medium, era, texture, line quality, color palette, and simplification level.
- Preserve subject and layout separately from style when using references.
- Avoid broad style words alone; describe visual traits such as flat color blocks, ink outlines, risograph grain, cel shading, or paper texture.
Composition Controls
- Placement: centered, lower third, top-left, rule of thirds, symmetrical, full-bleed, isolated on white, copy space on the left.
- Camera: eye-level, top-down, three-quarter view, macro, wide shot, telephoto compression, orthographic, isometric.
- Crop: close-up, waist-up, full-body, uncropped product, margin-safe, bleed-safe, no object cut off.
- Background: seamless studio, transparent, simple gradient, shallow-focus environment, contextual scene, empty copy space.
- Overlap: fully visible, not cropped, no objects covering labels, foreground partly obscures background, label unobstructed.
- Platform: square icon, vertical story, horizontal banner, wide hero, app-store screenshot, e-commerce listing.
Prompt Rewriting Workflow
1. Identify model, mode, source images, masks, aspect ratio, quality target, and final use. 2. Choose the matching task pattern from this guide. 3. Start with the desired deliverable and the most important subject attributes. 4. Add composition, background, lighting, palette, and style. 5. Add exact text in quotes or state “no text.” 6. For edits and references, put preservation and reference roles before creative changes. 7. End with avoid-list constraints and move operational settings into Vofy flags. 8. If the request is complex, split into draft, edit, and final-quality passes.
Examples
E-Commerce Hero
Create a premium e-commerce hero image of black wireless earbuds in an open charging case, three-quarter macro view, centered with 40% empty space on the left for copy. Matte charcoal seamless background, soft rim light on the case edge, subtle reflection, crisp commercial product photography. No text.Educational Diagram
Create a clean educational diagram showing the water cycle for middle-school students. Use four large labeled stages: "Evaporation", "Condensation", "Precipitation", "Collection". Blue arrows move clockwise through the diagram. Flat vector style, white background, high contrast, large readable labels. No extra text.Reference-Based Edit
Preserve the chair shape, wood grain, camera angle, and shadow from the reference image. Change only the seat cushion fabric to deep navy velvet. Match the original perspective, seams, edge softness, and studio lighting. Do not alter the chair legs or background.Transparent Asset
Create a transparent-background PNG-style asset of a glossy red heart-shaped balloon, front-facing with a tied knot and short curled ribbon. Smooth clean alpha edges, subtle internal highlight, realistic latex reflections, no backdrop, no ground shadow, no text.Product Mockup
Use image 1 as the exact product reference. Preserve the bottle silhouette, cap shape, front label layout, logo placement, and visible text. Create a premium studio mockup on warm beige paper, three-quarter front view, softbox reflection from upper left, subtle contact shadow, accurate glass thickness. Do not redesign the label or change the brand colors.Text Poster
Create a minimalist concert poster on warm off-white paper. Large title at top: "MOON ROOM". Smaller subtitle below: "LIVE 9 PM". Centered blue crescent moon icon, generous whitespace, modern Swiss grid, crisp black typography, no extra text.UI Mockup
Create a clean mobile banking app dashboard mockup on a single phone screen. Top greeting says "Good morning". Main card shows "$2,480" in large readable numbers. Three rounded action buttons: "Send", "Save", "Pay". Minimal white-and-navy interface, consistent spacing, soft shadows, no extra labels.Style Transfer
Use the reference image only for the subject and pose. Recreate it as a 1970s Japanese travel poster: flat color blocks, sun-faded paper texture, simplified shadows, teal-orange palette, clean border. Preserve the subject silhouette and direction of gaze. No text.Common Failure Fixes
- Too much creative freedom: add deliverable, use case, composition, lighting, palette, and avoid constraints.
- Composition mismatch: specify viewpoint, crop, object placement, whitespace, and background.
- Text errors: shorten text, quote exact words, increase size, simplify layout, and request no extra text.
- Reference drift: identify the primary reference and explicitly state what to preserve and ignore.
- Unnatural edit: ask to match shadows, reflections, perspective, grain, edge softness, and material.
- Weak transparent cutout: request clean alpha edges, isolated subject, no backdrop, no halo, and controlled shadow behavior.
- Product redesign: repeat preserve rules for silhouette, label layout, logo placement, colors, and visible text.
- Overloaded request: split into separate generation, local edit, and final-quality passes.
- Costly iteration: draft at lower quality, then raise quality after composition, text, and references are approved.
Final Rewrite Checklist
- The prompt states the asset type, use case, subject, and mode.
- Composition, camera, background, style, lighting, and color are concrete.
- Text is quoted, short, placed, and hierarchy-aware.
- Reference roles, primary reference, and ignore rules are explicit.
- Edit scope is local, preservation comes first, and integration details are included.
- Product, brand, UI, and character identity constraints are protected when relevant.
- Transparency, diagrams, and icons include readability or edge-quality constraints.
- Quality/latency/cost tradeoff is reflected in parameter hints.
- Operational settings are Vofy flags, not prompt prose.
Sora 2 Prompt Guide
Source: https://developers.openai.com/cookbook/examples/sora/sora2_prompting_guide
Applies to: sora-2, sora-2-pro.
Use this as an agent reference for rewriting prompts for Vofy/Sora 2. It is intentionally fuller than SKILL.md: keep workflow instructions short there and keep model-specific prompting detail here.
What Sora 2 Can Be Steered With
- Text prompts can control video style, setting, subject, action, camera movement, dialogue, ambient sound, and continuity.
- Image input can anchor the first frame, composition, character appearance, product design, wardrobe, environment, and aesthetic.
- Character references can improve consistency across people, animals, or objects when supported by the route.
- Edits and extensions work best when the prompt says what stays fixed before describing what changes.
- Sora 2 responds well to cinematic language, but the strongest instructions are still visible, audible, and physically concrete.
API And Vofy Boundaries
Do not bury hard parameters in prose. Prefer Vofy flags and capability docs:
--model:sora-2orsora-2-pro.--duration: Sora API supports4,8,12,16, and20seconds; verify current Vofy support withvofy models <model>.--aspect-ratio/--resolution: choose through supported model constraints, not prompt text alone.sora-2: 720p portrait or landscape (720x1280,1280x720) when exposed by Vofy.sora-2-pro: 720p, 1024p, and 1080p portrait/landscape when exposed by Vofy.- Quality tradeoff: higher resolution can improve detail and consistency but may take longer and cost more.
- Endpoints/modes differ: generation, remix/edit, extension, and character-reference workflows may expose different inputs.
- Use Vofy input flags for first-frame media, reference images/videos, remixes, extensions, and character IDs; do not invent unsupported flags.
Prompting Mindset
- Write like a director or cinematographer briefing a shoot: what is on screen, how it moves, how it sounds, and how it is lit.
- Keep the first sentence high-signal: output style, main subject, setting, and main action.
- Choose the amount of detail intentionally: detailed prompts improve control; lighter prompts allow creative variation.
- Avoid overloading one shot with many simultaneous events, competing subjects, or contradictory styles.
- Iterate systematically. Change one variable at a time: camera, lens, action, lighting, color, dialogue, or continuity anchor.
- Short clips are easier to control. Split complex concepts into multiple clips or shot blocks.
Detail Levels
Lightweight Prompt
Use when creative variation is welcome.
A handheld documentary-style shot of a street musician playing violin under a subway entrance at night. Rain glints on the pavement, commuters pass behind him, and the camera slowly pushes in as the final note rings out.Descriptive Prompt
Use for most agent rewrites.
Style: [documentary / commercial / 16mm / animation / phone footage], [texture/grade].
Scene: [subject, wardrobe/materials, setting, weather, props].
Cinematography:
Camera shot: [framing and angle].
Camera motion: [locked-off / slow push-in / handheld follow / orbit / tilt].
Lens/depth: [wide deep focus / macro shallow focus / telephoto compression / anamorphic flare].
Lighting + palette: [source direction and quality], [3-5 color anchors].
Mood: [tone].
Actions:
- [beat 1: clear visible gesture or movement]
- [beat 2: next chronological beat]
- [final beat: ending pose, hold, or transition]
Dialogue:
- [Speaker, action]: "[short natural line]"
Background sound: [diegetic ambience only, or silence].
Avoid: [only critical exclusions].Ultra-Detailed Prompt
Use for strict cinematography, VFX, product, or continuity needs.
Format & look: [capture medium], [grain/halation/texture], [grade].
Location & framing: [foreground], [midground], [background], [shot scale], [angle].
Subject/wardrobe/props: [stable identity anchors], [materials], [required text/logos].
Lighting & atmosphere: [key/fill/rim/practicals], [weather/haze/particles].
Lenses & filtration: [focal length feel], [depth of field], [filter artifacts].
Shot timing:
- 0.00-[time]s: [beat]
- [time]-[time]s: [beat]
- final second: [hold or payoff]
Sound: [diegetic audio, dialogue, no score if needed].
Continuity notes: [what must remain stable across shots/edits/extensions].Core Prompt Anatomy
[Format/style], [shot scale and angle] of [subject with identity anchors] in [setting].
Start: [initial visible state].
Action: [chronological beats that fit the duration].
Camera: [one setup and one movement], [lens/framing/depth-of-field feel].
Lighting/color: [specific sources, mood, 3-5 palette anchors].
Audio: [brief dialogue, ambience, or silence cue].
Continuity: preserve [identity, outfit/product, geography, lighting, motion direction].
End: [final state or held frame].Use only fields that matter. Concrete nouns, visible actions, and ordered timing beat generic adjectives.
Subject And Scene Control
- Introduce the main subject early with stable identity anchors: age range, silhouette, outfit, material, color, logo/text, or distinctive object features.
- Describe the environment as layered composition: foreground, midground, background.
- Use specific props and textures: cracked ceramic mug, chrome espresso machine, fogged glass, frayed denim, paper lanterns.
- Make spatial relationships explicit: “the bicycle is leaning against the left wall,” “the red suitcase stays beside her right foot.”
- For products, specify exact visible surfaces, required labels, packaging shape, and what must remain legible.
- For animals or characters, avoid unnecessary trait changes between sentences.
Style And Medium
Good style descriptions combine medium, era, capture feel, and texture:
- Documentary realism: handheld, available light, imperfect framing, natural background action.
- Commercial product: glossy highlights, clean negative space, controlled reflections, macro detail.
- Film look: 16mm grain, halation, gate weave, muted stock colors, soft highlight rolloff.
- Smartphone footage: vertical framing, casual hand shake, compressed dynamic range, autofocus breathing.
- Animation: claymation, watercolor, cel animation, stop motion, low-poly, miniature diorama.
- VFX/fantasy: practical set feel, believable scale cues, contact shadows, particles, atmospheric depth.
Avoid vague style-only prompts like cinematic, beautiful, or epic unless paired with concrete visual choices.
Camera And Composition
- Pick one camera setup per shot unless writing a deliberate multi-shot sequence.
- Specify shot scale: extreme wide, wide, medium, close-up, macro, over-the-shoulder, POV, aerial.
- Specify angle: eye-level, low-angle, high-angle, top-down, Dutch tilt, profile, three-quarter view.
- Specify motion: locked-off, slow push-in, dolly out, pan left, tilt down, handheld follow, orbit, crane up.
- Use one lens/depth cue: macro shallow focus, deep-focus wide shot, telephoto compression, anamorphic flare.
- If the output drifts, use locked camera, simple background, and fewer moving subjects.
Lighting, Color, And Atmosphere
- Name light sources: window key, tungsten practical lamp, neon sign, moon rim, overcast skylight, car headlights.
- State direction and quality: soft left key, hard backlight, low warm practicals, cool overhead fluorescents.
- Use palette anchors instead of broad adjectives: amber, teal, cream, oxblood, graphite, sodium orange.
- Include atmosphere only when important: haze, dust motes, rain mist, smoke plume, underwater particles.
- Keep lighting continuity explicit in edits/extensions: preserve source direction, time of day, shadow length, color temperature.
- For consistent series, repeat palette and light-source anchors across prompts.
Motion And Timing
- One primary subject action plus one primary camera move is the most reliable pattern.
- Put actions in chronological order. Avoid many simultaneous unrelated actions.
- Use physical verbs and counts: steps, turns, lifts, pours, blinks, brakes, opens, folds, lands.
- Tie action density to duration:
- 4 seconds: one clear movement plus an ending hold.
- 8 seconds: 2-3 beats with simple transitions.
- 12-20 seconds: structured shot blocks or a small scene arc.
- Use final-second payoffs: pause, held expression, object reveal, completed turn, settled frame.
- If motion becomes chaotic, reduce moving background elements and make the camera locked-off.
Dialogue And Audio
- Use a labeled
Dialogue:block for speech. - Keep lines short enough to be spoken naturally during the clip.
- Attribute every line to a speaker and pair it with visible face/body action.
- For 4 seconds, use one or two short lines; for 8 seconds, a few brief turns can work.
- Background sound should support the scene: rain on glass, rail brakes, espresso machine hum, distant traffic, crowd murmur.
- Avoid asking for a full mixed soundtrack when diegetic ambience is enough.
- If a Vofy route does not support audio, convert speech into visible acting or omit audio.
Dialogue:
- Detective, whispering: "You're lying."
- Suspect, looking away: "Maybe I'm just tired."
Background sound: rain on the window and a distant police radio.Image Input And First Frames
- Use image input to anchor the first frame, composition, character design, wardrobe, set dressing, product layout, or aesthetic.
- Preserve first, animate second: identity, outfit/product design, composition, lighting, and text/logo placement come before motion.
- For first-frame prompts, describe what happens after the provided frame rather than redescribing the entire image.
- Ensure the input image matches target video resolution when the route requires it.
- Animate only a few attributes: expression, hand motion, fabric movement, camera drift, environmental motion.
- Mention important static elements that must not change: logo placement, face, outfit, prop geometry, room layout.
Use the provided first frame as the opening frame. Preserve the woman’s face, red coat, rainy window lighting, and camera angle. She slowly turns toward the glass, exhales a small fog patch, then smiles faintly. Camera remains locked. End on the same composition.Character References
- Use character references when supported to keep a person, animal, or object consistent across generations.
- Character references are created from reference video assets outside the prompt; prompts should use the actual character handle/ID supplied by Vofy.
- Short, clear reference clips work best: isolated subject, stable lighting, visible defining features, minimal occlusion, limited background distractions.
- Reference no more than two characters per generation when possible.
- In the prompt, use the same name throughout and repeat only essential anchors: face, build, outfit, species/breed, signature accessory.
- Avoid introducing conflicting age, hairstyle, color, size, or wardrobe details after the character has been anchored.
A cinematic shot of Alfie running through wet grass at sunrise. Preserve Alfie’s small terrier build, cream curls, red collar, and playful ears. Camera follows low behind him as he bounds three times, turns toward camera, and pauses in warm rim light.Remix And Edits
- Start with what stays unchanged, then state the edit.
- Change one variable at a time: color grade, lens, prop, action, background detail, wardrobe, added subject, or weather.
- Preserve continuity anchors: identity, pose, camera angle, composition, light direction, product labels, and scene geography.
- Avoid edits that contradict the source frame or require large hidden geometry changes unless that is the purpose.
- When close to the target, pin successful details before requesting the next tweak.
Preserve the original camera angle, woman’s pose, black dress, and candlelit restaurant background. Change only the weather visible through the window: heavy rain streaks down the glass with occasional blue lightning flashes. Keep the warm indoor lighting unchanged.Extensions
- Extensions use the full original clip as context; state continuity first, then the next action.
- Preserve camera direction, lens feel, lighting, character pose, soundbed, motion direction, and scene geography.
- Add the next logical beat rather than summarizing the original clip.
- Individual extensions can be up to 20 seconds in the Sora API and can be chained when supported by the client.
- For chained extensions, keep a continuity log: final pose, camera direction, lighting, active props, and audio bed.
Continue from the final frame. Preserve the locked-off camera, warm desk-lamp lighting, rain sound, and the character’s seated posture. She closes the notebook, looks toward the door, and the lamp flickers once before the shot holds on her concerned expression.Multi-Shot Prompts
- Multi-shot prompts can work, but each block needs a clear boundary.
- Give each shot one camera setup, one main action, and one lighting recipe.
- Prefer separate generations when exact editing rhythm matters.
- Use shot blocks when continuity matters more than frame-perfect cutting.
Shot 1, 0-4s: Wide eye-level establishing shot of the empty rooftop at golden hour; sheets sway and city traffic hums below.
Shot 2, 4-8s: Medium close-up of the dancer entering frame; slow dolly-in, warm edge light, red silk dress catching wind.Negative Instructions And Avoid Lists
- Use
Avoid:sparingly for critical exclusions only. - Prefer positive constraints over long negative lists: “locked camera, empty background” is clearer than many “no...” phrases.
- Good avoid items: extra fingers, text changes, logo changes, camera cuts, additional people, warped product shape, non-diegetic music.
- Do not include broad contradictory avoid lists that compete with the main prompt.
Troubleshooting Rewrite Patterns
- Identity drift: add stable anchors, use character/reference input, reduce action, repeat outfit and face constraints.
- Camera drift: specify locked-off tripod or one simple movement; remove competing camera language.
- Chaotic motion: reduce background actors, split into beats, shorten duration, use final hold.
- Weak product accuracy: describe shape, material, label placement, front-facing angle, and what text must remain unchanged.
- Bad dialogue timing: shorten lines, reduce speakers, pair speech with visible mouth/face action.
- Inconsistent lighting: name source direction, time of day, palette, and shadows; repeat them in extension/edit prompts.
- Overconstrained result: remove secondary props, redundant adjectives, and low-priority avoid items.
Weak To Strong Transformations
- Weak:
A beautiful street at night. - Strong:
Wet asphalt, zebra crosswalk, neon signs reflected in puddles, and a single cyclist braking at the curb under blue-pink storefront light.
- Weak:
Person moves quickly. - Strong:
The cyclist pedals three times, brakes hard, and stops with one foot down in the final second.
- Weak:
Cinematic look. - Strong:
Wide low-angle shot, shallow depth of field, anamorphic lens feel, warm rim light through fog, amber and teal palette.
- Weak:
Make this image move. - Strong:
Use the provided image as the opening frame. Preserve the product position, label, marble counter, and soft left window light. Only animate condensation sliding down the bottle and a subtle camera push-in.
- Weak:
Continue the video. - Strong:
Continue from the final frame. Preserve the handheld forward motion, blue dusk lighting, wet street reflections, and distant siren sound. The runner slows, looks back once, then turns into the alley as the camera follows.
Complete Example: Product Clip
Style: glossy product commercial with clean macro detail and soft studio reflections.
Scene: a matte black ceramic coffee mug on a walnut desk beside cream paper and a silver spoon. The mug logo faces camera and stays readable.
Cinematography:
Camera shot: centered macro close-up at desk height.
Camera motion: slow push-in only.
Lens/depth: shallow depth of field, crisp logo, soft background falloff.
Lighting + palette: large softbox from upper left, warm rim from a tungsten desk lamp; black, walnut brown, cream, and amber.
Actions:
- Steam curls upward from the mug in thin strands.
- A hand enters from the right and gently places the spoon beside the mug.
- Final second holds on the readable logo and rising steam.
Background sound: quiet room tone and a faint ceramic clink.
Avoid: changing the logo text, extra hands, camera cuts.Complete Example: Character Clip
Style: naturalistic handheld documentary footage, early morning park.
Scene: Maya, wearing a yellow raincoat and white sneakers, stands on a wet path under maple trees. Fallen orange leaves cover the ground.
Cinematography:
Camera shot: medium eye-level shot from three meters away.
Camera motion: gentle handheld follow as she walks toward camera.
Lens/depth: natural smartphone-like depth, background slightly soft.
Lighting + palette: overcast skylight, wet green leaves, yellow coat, gray path, orange leaves.
Actions:
- Maya looks down, steps around a puddle, and laughs softly.
- She raises one hand to catch a falling leaf.
- Final second holds as she looks into camera with the leaf in her palm.
Dialogue:
- Maya, smiling: "I found the perfect one."
Background sound: light rain on leaves and distant city traffic.
Continuity: preserve Maya’s yellow raincoat, white sneakers, path direction, and overcast lighting.Complete Example: Extension
Continue from the final frame of the provided clip. Preserve the same handheld camera height, forward walking direction, wet pavement reflections, blue-pink neon palette, and distant traffic sound. The cyclist pushes off from the curb, pedals twice through the crosswalk, then brakes under the next streetlight. Camera follows from behind at walking speed and holds as the red brake light reflects in the puddle.Final Rewrite Checklist
- Subject, setting, style, and action are clear in the first sentence.
- Prompt detail level matches the goal: creative variation or tight control.
- Hard parameters are represented as Vofy flags, not only prompt prose.
- Motion is chronological, physically plausible, and duration-aware.
- Camera has one explicit setup and movement per shot.
- Lighting, palette, texture, and depth of field are concrete.
- Dialogue/audio is short, labeled, naturally timed, and route-supported.
- Reference preservation comes before animation, edits, or extensions.
- Character names, handles, and identity anchors stay consistent.
- Avoid list is short and critical, not a competing prompt.
Veo Prompt Guide
Source: https://ai.google.dev/gemini-api/docs/video?hl=zh-cn&example=dialogue#prompt-guide
Applies to: veo-3.1, veo-3.1-fast, veo-3.1-lite.
Model Behavior
- Veo prompts should read like a short video brief, not a keyword pile.
- Start with the core idea, then add concrete video terms and modifiers.
- Specify subject, context, action, style, camera motion, composition, and ambiance.
- Add lighting, color, sound, lens, and focus details when they affect the result.
- Motion reliability improves when actions are chronological and duration-aware.
- Dialogue, sound effects, and ambient sound can be prompted when the selected Vofy model, mode, and flags support audio.
- Use Vofy flags for duration, aspect ratio, resolution, audio, first/last frames, references, and modes.
- Do not prompt for harmful, unsafe, or policy-blocked content; Veo may reject prompts or outputs.
Core Prompt Anatomy
[Style/format], [shot type] of [subject] [main action] in [location/context].
Start: [initial state].
Action: [ordered motion/change that fits duration].
Camera: [movement, angle, framing].
Composition: [foreground/midground/background or subject placement].
Lens/focus: [depth of field, macro, wide-angle, soft focus, focal emphasis].
Ambiance: [lighting, color palette, mood, weather, environmental sound].
End: [final state].Use the full anatomy for complex video prompts. For simple prompts, keep one fluent paragraph with the same information.
Prompt Elements
- Subject: person, animal, product, vehicle, object, environment, or character identity.
- Context: location, time of day, weather, props, crowd level, architecture, background activity.
- Action: one clear visible action, or 2-3 simple beats for longer clips.
- Style: cinematic, documentary, commercial, handheld phone, animated, stop-motion, hyperreal, macro, slow motion.
- Camera motion: dolly in, pan left, tracking shot, orbit, crane up, locked-off tripod, handheld follow.
- Composition: wide shot, medium close-up, close-up, low angle, overhead, centered, over-the-shoulder.
- Focus and lens effects: shallow depth of field, deep focus, rack focus, soft focus, macro, wide-angle, reflected details.
- Ambiance: color scheme, lighting, mood, weather, environmental motion, sound effects, ambient sound.
Use descriptive adjectives and adverbs, but keep each detail tied to something visible or audible in the clip.
Prompt Modifiers
Add modifiers only when they improve control:
- Lighting: golden-hour backlight, soft studio reflections, neon glow, moonlit blue shadows, overcast diffuse light.
- Color: muted earth tones, high-contrast black and gold, cool blue palette, saturated candy colors.
- Pacing: slow graceful movement, energetic handheld pace, subtle idle motion, deliberate reveal.
- Texture: glossy metal, handmade paper, misty glass, dusty shelves, wet asphalt.
- Atmosphere: quiet room tone, distant traffic, ocean wind, cafe murmur, mechanical click.
Avoid stacking many unrelated modifiers. Prioritize the 3-5 traits that define the shot.
Duration-Aware Beat Planning
For 4-6 seconds:
Start: [initial pose/scene].
Action: [one visible movement].
End: [clear final pose/object state].For 8+ seconds:
0-2s: [establish subject and setting].
2-6s: [main action].
6-8s+: [reaction, reveal, or ending hold].Do not request many locations, fast cuts, or complex story arcs in a short clip. If the story needs multiple shots, generate separate clips.
Dialogue And Audio
- Keep spoken lines short and natural.
- Attribute each line to a speaker.
- Put exact spoken lines in quotation marks.
- Match line length to duration.
- Describe facial expression or action around the line.
- Add sound effects and ambient noise explicitly when they matter.
- If audio is unsupported, convert dialogue/sound into visual acting or omit it.
Medium close-up dialogue scene in a rainy cafe. A woman looks up from a notebook, half-smiles, and says, "I think we found it." Camera slowly pushes in, warm cafe lights behind her, raindrops streaking the glass, quiet room tone.For face-forward clips, specify portrait framing, expression, eye direction, and the facial detail that should remain in focus.
Camera Guidance
- Use one primary camera move per prompt.
- Combine shot scale and motion: medium close-up with slow push-in, wide aerial tracking shot, low-angle locked-off shot.
- If action is complex, keep camera still.
- If camera is dynamic, keep subject action simple.
- Avoid contradictory instructions like “locked-off handheld orbit.”
- Avoid vague camera phrases like “make it cinematic” unless paired with specific framing, light, and motion.
Image-To-Video
Start with preservation:
Use the provided image as the starting frame. Preserve [subject identity, outfit/product design, composition, lighting, text/logo].
Animate only [specific motion]. Camera [movement or locked]. End with [final state].Choose an input image that already resembles the desired opening frame. For drawings, paintings, products, or natural scenes, describe the motion and sound to add rather than redescribing every visible detail.
For product animation, preserve product geometry, logo placement, and label text; animate camera, lighting, environment, or simple object motion.
First And Last Frame
Use first + last frame mode when the beginning and ending states matter more than the exact path.
Use the first image as the opening frame and the second image as the final frame. Preserve [subject/style/setting]. Create a smooth, plausible transition where [subject/action/change] connects the two frames. Camera [movement or locked].- Keep the transition physically plausible.
- Avoid asking for new subjects or settings that are not present in either frame.
- Use simple motion when the two frames differ significantly.
Reference Assets
- Assign each reference a role: identity, style, movement, environment, product design, or sound.
- For Veo 3.1 reference images, use up to three asset references for the same person, character, or product when appearance preservation matters.
- Use reference images for consistency, not as a substitute for describing the desired action.
- If references conflict, state which reference wins for identity, style, or setting.
Use reference image 1 for the character identity and outfit. Use reference image 2 for the neon city lighting style. Keep the character's face, hair, and jacket consistent while she walks through a rain-soaked alley.Video Extension
- Continue from the final second of the clip.
- Preserve existing subject state, camera direction, lighting, motion speed, and scene geography.
- Add only the next natural beat instead of restarting the scene.
- If the final second is silent, do not rely on strong audio continuation.
Continue from the final moment of the provided video. Preserve the same handheld camera direction, rainy night lighting, and walking pace. The character turns toward a glowing storefront, pauses, and reaches for the door handle as traffic reflections ripple on the pavement.Aspect Ratio And Platform Fit
Set aspect ratio with Vofy flags when supported; reinforce composition in the prompt only when the framing matters.
- 16:9: cinematic landscape, wide environments, product reveals, group action.
- 9:16: mobile social video, portrait framing, single subject, vertical motion.
- 1:1: centered product, icon-like compositions, compact social previews.
Do not rely on prompt text alone for aspect ratio, duration, or resolution when Vofy exposes flags.
Negative Guidance
Veo prompts usually work better with positive direction than long negative lists.
Prefer:
Keep the camera locked, the subject centered, and the background softly blurred.Instead of:
No shaky camera, no extra people, no background clutter, no blur problems.Use explicit “do not change” constraints for source-driven modes when preservation is critical.
Style Examples
Cinematic Scene
Cinematic wide shot of a lone hiker crossing a black volcanic beach at dawn. The scene starts with the hiker small against the shoreline, then she walks toward a thin beam of golden light breaking through clouds. Camera slowly tracks sideways, waves rolling in the background, cool blue shadows with warm sunrise rim light, quiet wind ambiance.Product Commercial
Premium macro commercial shot of a silver smartwatch on a dark stone surface. The clip starts on a close-up of the crown, then the camera slowly orbits to reveal the screen lighting up. Soft studio reflections glide across the metal edge, black and cool-blue palette, subtle mechanical click sound if audio is supported.Dialogue
Medium close-up dialogue scene in a small bookstore at night. A young bookseller closes an old ledger, looks toward someone off-camera, and says, "This copy was never meant to be sold." Camera slowly pushes in, warm lamp light, dusty shelves behind her, quiet street rain outside.Animated Style
Whimsical stop-motion style shot of a paper boat sailing across a kitchen sink like an ocean. The boat bobs through tiny soap bubbles, passes a spoon like a silver island, and ends under a dripping faucet. Camera is locked-off, warm afternoon light, handmade paper texture.Portrait With Audio
Vertical medium close-up of a chef in a bright test kitchen, framed from chest up. She smiles directly at camera and says, "The secret is patience." Camera is locked, shallow depth of field keeps her eyes sharp, warm overhead light, soft kitchen clatter in the background.Natural Scene From Image
Use the provided forest image as the starting frame. Preserve the mossy rocks, tall pine trees, and misty morning light. Animate a gentle breeze moving the branches while thin fog drifts between the trunks. Camera slowly pushes forward along the path, quiet bird calls and distant water if audio is supported.Common Failure Fixes
- Too much happens: reduce to one subject, one action, one camera move.
- Weak motion: add start, middle, and end beats.
- Camera confusion: choose either moving camera or locked camera.
- Style mismatch: add style, lighting, palette, texture, and medium.
- Flat composition: specify shot scale, subject placement, foreground/background, and focus.
- Dialogue problems: shorten the line, attribute the speaker, and describe expression.
- Audio mismatch: name the exact sound source and keep it consistent with the scene.
- Reference drift: start with preservation instructions and assign each reference a role.
- Product/logo drift: explicitly preserve geometry, logo placement, label text, colors, and materials.
- Unrealistic transition: simplify motion between first and last frames.
Final Rewrite Checklist
- Subject, action, and setting are explicit.
- Prompt uses a natural video brief, not only keyword tags.
- Motion is chronological and duration-aware.
- Camera movement and shot scale are specified.
- Composition and subject placement are clear.
- Focus, lens, lighting, palette, and ambiance are concrete.
- Dialogue/audio is short, attributed, and supported by the chosen model/mode.
- Reference preservation comes before animation instructions.
- Aspect ratio, duration, resolution, and mode are handled with Vofy flags when available.
- Prompt avoids unsafe content and unsupported controls.