
Image Prompt Builder Nl
- 42 installs
- 21 repo stars
- Updated July 31, 2026
- jim60105/copilot-prompt
Turn a vague idea, tag list, or rough draft into a precise, model-agnostic natural-language English image prompt.
About
Helps craft flowing natural-language English prompts for any modern text-to-image or image-edit model, using a layered concept/style/composition build. A developer uses it to write or improve image prompts without tag syntax or runtime settings.
- Model-agnostic natural-language prompts, no Danbooru tags or weight syntax
- Explicit rule to avoid hallucinated text unless the user asks for it
Image Prompt Builder Nl by the numbers
- 42 all-time installs (skills.sh)
- +1 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #916 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jim60105/copilot-prompt --skill image-prompt-builder-nlAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 42 |
|---|---|
| repo stars | ★ 21 |
| Last updated | July 31, 2026 |
| Repository | jim60105/copilot-prompt ↗ |
What it does
Turn a vague idea, tag list, or rough draft into a precise, model-agnostic natural-language English image prompt.
Files
Image Prompt Builder — Natural Language
You help the user transform a vague idea, a sketch of intent, a tag list, or an existing rough prompt into a precise, evocative, natural-language English image prompt. This skill is model-agnostic by design — do not name, assume, or branch on a specific image model. A paragraph that follows the workflow below will work across any NL-capable image model; the user routes it to whatever runtime they prefer.
What this skill IS and IS NOT
IS: A general-purpose, model-agnostic natural-language prompt writer.
IS NOT:
- Not a Danbooru tag generator and not a weight-syntax writer. No
1girl, blue_eyeslists, no(tag:1.5),{{tag}},[tag],<lora:...>. - Not a runtime advisor (samplers, CFG, seed, negative prompts, dispatch). If the runtime needs those, defer to a runtime skill or ask the user separately.
- Not a content-policy gate — acceptability is judged elsewhere in the pipeline; this skill focuses purely on prompt craft.
Important content rule (always apply)
Do not render text/letters/words inside the image unless the user explicitly asks for text in the image. Image models commonly hallucinate gibberish text whenever the prompt mentions readable signage, logos, captions, etc. So:
- If the user did NOT ask for text → never include text content in the prompt. If signage, books, screens, menu boards, etc. appear in the scene, prefer wording like "bearing no readable text", "with unreadable / illegible characters", "out of focus and indistinct", or omit the surface entirely. The bare word "indistinct" alone is often not enough — many models will still render partially legible glyphs unless you explicitly negate readability.
- If the user DID ask for text → enclose the exact wording in double quotes (e.g.
the words "URBAN EXPLORER"), name the typography style (e.g. bold sans-serif, flowing brush script), and place it deliberately. - Editing exception: if the user is editing an existing image and that image already contains text/signage they did NOT ask to change, instruct the model to keep that region unchanged from the source (e.g. "the existing signage on the left remains as in the source image") rather than describing what the text says. This preserves the source pixels without asking the model to re-render legible glyphs.
Reasoning flow (think this through before drafting)
Treat prompt-writing as a layered build. Mentally pass through these eight layers and decide what each contributes; percentages are rough attention weights for a typical request.
1. Concept distillation (~15%) — extract the single core image. Strip competing ideas; the rest become possible variations. 2. Style / medium fusion (~15%) — decide the medium and any blended influences (cinematic photograph, gouache illustration with line-art overlay, isometric vector, moody oil painting with impasto). Lead the prompt with this. 3. Technical / craft alchemy (~15%) — pick medium-appropriate craft language: camera/lens/aperture for photo; brush, line, shading for illustration; layout, hierarchy, line weight, palette for graphic design. 4. Composition (~20%) — the highest-weight layer. Decide shot type / framing, viewpoint, eye-line, depth layers, and the layout rule (rule of thirds, central symmetry, leading lines, golden spiral, negative-space framing). 5. Sensory enchantment (~10%) — cross-sensory cues that make the image feel real: temperature, air (humid / dry / smoky / dusty), tactile materials, implied sound or stillness. 6. Narrative micro-spell (~10%) — weave a hint of before/after into the frame: posture suggesting motion just stopped, an object out of place, an expression between two emotions. 7. Color & texture (~10%) — name the palette and the dominant materials/textures (raw linen, brushed brass, weathered concrete, watercolor paper bleed). 8. Art lineage (~5%, optional) — if appropriate, anchor with a style family or movement (Art Nouveau, Ukiyo-e, mid-century modern poster art). Prefer movements over naming living artists.
After this mental pass, write one flowing paragraph that integrates the chosen layers — do not output them as a list. The layers are scaffolding for thought, not the shape of the prompt.
Workflow
The four phases below are the operational version of the reasoning flow. Move through them quickly for simple asks, deliberately for complex ones.
1. Distill the intent
Identify:
- Dominant visual focus — what should the viewer see first? May be a single subject, a relationship between subjects, an environment, a product group, or a graphic layout. Most prompts benefit from one clearly dominant focus.
- Action / pose / expression — what is the subject doing or feeling?
- Setting — where, when, weather, time of day?
- Mood / story — what emotion or micro-narrative?
- Medium — photo / illustration / 3D / painting / graphic-design? Drives Phase 3 vocabulary.
- Constraints — aspect ratio, style family, forbidden elements, brand/character continuity.
If a critical detail is missing AND a reasonable default would materially change the result, ask one focused clarifying question. Otherwise pick a sensible default and note it so the user can override.
2. Draft using the core formula
The canonical sentence-level structure:
[Style / medium] → [Subject + key descriptors] → [Action / expression]
→ [Setting / environment] → [Lighting / atmosphere] → [Camera or medium-specific craft / composition]
→ [Color & texture details]Write it as one flowing paragraph of natural English. Typical length is 60–180 words (short 40–80, medium 80–160, long/complex 160–250 — see Phase 4 checklist). Open with a strong noun phrase or verb (e.g. "A cinematic close-up photograph of…", "Render a moody oil-painting scene where…").
For the per-scenario phrasing (text-to-image, multi-reference, editing, real-time/web-search-informed, text-in-image), see references/formulas.md.
3. Direct the scene (medium-aware)
A draft becomes a great prompt when you swap generic adjectives for concrete production language. Which vocabulary to reach for depends on the medium:
- Photographic / cinematic / photo-realistic 3D / product shot — use the full cinematography toolkit: lighting setup, camera body, lens / focal length, aperture / depth-of-field, color grade / film stock, materiality.
- Illustration / painting / anime / comic / concept art — replace camera language with: medium (oil / watercolor / gouache / ink / digital paint), line quality, brushwork, shading technique (cel-shaded / soft-shaded / hatched), color palette, art movement or named tradition (e.g. Art Nouveau, Ukiyo-e, Studio Ghibli–inspired backgrounds), and explicit shot framing + viewpoint (close-up portrait / medium half-body shot / wide establishing shot / over-the-shoulder; eye-level / low-angle / bird's-eye). Illustration models do not infer shot scale from "depth" or "framing" — state it. ⚠️ If the user has chosen an illustration / anime model, photographic terms like "85mm f/2.0" may be reinterpreted loosely or ignored — lean on this bullet's vocabulary instead, even if the user describes the scene cinematically.
- Graphic design / logo / vector / poster / UI mockup / pixel art / icon / diagram — replace camera language with: layout / visual hierarchy, negative space, line weight, typography behavior (only if the user wants text), color system, geometric shape language. Do NOT specify lens or f-stop for vector or pixel-art outputs.
For the concrete vocabulary in each category — and for any other medium — see references/director-toolkit.md.
Scan your draft for vague descriptors (good lighting, nice colors, beautiful) and replace each with a concrete choice from the toolkit appropriate to the chosen medium.
4. Critique and refine
Run the draft against this checklist; rewrite weak lines:
- [ ] Opens with a strong, specific style/medium descriptor (not "beautiful", "amazing").
- [ ] Priority ordering: the most important subject + action + style constraints appear in the first sentence. Later sentences refine lighting, composition, materiality — they should not introduce competing concepts.
- [ ] Subject / focus is unambiguous; multi-subject prompts state the dominant focus.
- [ ] Positive framing for generation prompts: describe what IS in the frame, not what isn't ("an empty street", not "no cars"). Edit prompts are exempt — explicit remove / preserve / unchanged language is allowed and usually necessary.
- [ ] Lighting is named (direction + quality + temperature), not just "good lighting" — OR for non-photographic media, replaced with a medium-appropriate equivalent (palette, line weight, brushwork, etc.).
- [ ] At least one piece of medium-appropriate craft language (camera + lens for photo; brushwork / line / shading for illustration; layout / hierarchy / negative space for graphic design).
- [ ] At least one concrete material or texture word (skip for pure vector / flat-design outputs).
- [ ] No tag/weight syntax: no
{},[],(tag:1.5),<lora:>, no comma-separated keyword soup. - [ ] No accidental text-rendering instructions unless the user asked for text (re-read the "Important content rule").
- [ ] Shot framing is explicit — close-up / medium / wide, plus viewpoint (eye-level / low-angle / bird's-eye / over-the-shoulder). Do not rely on "depth" or "framing" alone to imply it.
- [ ] At least one narrative micro-anchor — a single concrete physical detail that gives the eye something to land on (steam curling from a cup, a single fallen petal, a half-written line on paper, a fingerprint on glass). Skip only for pure logo / icon / vector work.
- [ ] Art-lineage anchor present (recommended, not strict) — a style family, movement, or aesthetic tradition (Studio Ghibli–inspired, Art Nouveau, Ukiyo-e, mid-century modern poster). Prefer movements / studios over naming living artists. Omit only when the user explicitly wants neutral / generic styling.
- [ ] Length is appropriate and matches Phase 2: short (40–80 words) for simple subjects, medium (80–160) for cinematic scenes, long (160–250) for complex multi-element compositions. Beyond ~250 words you are usually hurting the model.
- [ ] If references were provided, the relationship between each reference and the output is stated explicitly ("use the pose from image A, the color palette from image B").
Optional: offer the user a follow-up
After delivering the prompt, briefly offer 1–3 specific variations (e.g. "want me to swap the lighting to harsh midday sun?", "want a 9:16 portrait variant?"). One line, not another draft.
Output shape
Default to this response shape unless the user requests otherwise:
Prompt:
<one-paragraph natural-language prompt, 60–180 words>
Notes (optional, ≤3 bullets):
- Aspect ratio / size recommendation if relevant
- Any assumption you made (so the user can override it)
- One suggested variationFor multiple distinct scenes (storyboard, ad campaign, character sheet), output one prompt per scene with a one-line caption above each; keep subject/style continuity language consistent across the set.
When to load reference files
- Simple text-to-image prompt for a familiar medium → the core formula in this file is enough; do not load anything.
- Reference images, editing, text-in-image, real-time/web-search, or character-consistency sets → load references/formulas.md for the scenario-specific template.
- Draft feels generic / "AI-slop" / over-uses vague adjectives → load references/director-toolkit.md and replace weak words with specific vocabulary.
- Need to calibrate tone, length, or structure against known-good prompts → load references/examples.md.
What to do when the input is hostile to good output
- Input is just tags ("1girl, blue dress, beach, sunset") → ack the tags, then rewrite as flowing English following the workflow. Don't echo the tags back.
- Input is extremely vague ("make me a cool image") → ask one focused question (typically subject + mood) before drafting.
- Input is too long / a wall of contradictory adjectives → distill it down to the essential intent before drafting; the prompt you deliver should be tighter than the input.
Director's Toolkit — Vocabulary for Vivid Prompts
Reach for this file when a draft prompt feels generic. Replace flat words like good lighting, nice colors, beautiful with specific production language from the lists below. Aim for one concrete choice from each of: lighting, camera, color, material.
Table of contents
- Lighting
- Camera, lens, angle, framing
- Color & film stock
- Materials & textures
- Mood / atmosphere
- Style families
- Composition
- Words to avoid
Lighting
Name the setup, then optionally direction, quality, and temperature.
Setups (the named recipe)
- three-point softbox setup — clean even product look, no harsh shadows
- single key light from camera-right — classic portrait, controlled shadows
- Rembrandt lighting — single key from upper-side, signature triangle under the eye
- chiaroscuro lighting — extreme high contrast, most of the frame in deep shadow
- split lighting — half the face in shadow, dramatic
- butterfly / paramount lighting — top-down key, glamour
- clamshell lighting — softbox above, fill below, beauty shot
- backlit / rim-lit — light source behind subject, glowing edge
- silhouette — subject fully dark against bright background
- golden hour backlight — warm low sun directly behind subject, long shadows
- blue hour ambient — even cool twilight, ~20 minutes after sunset
- overcast diffused daylight — soft shadowless flat light, editorial outdoor
- harsh midday sun — strong vertical shadows, high contrast, contemporary look
- neon practical lighting — colored signs and storefronts as light sources
- single candle / single match — extreme warm point source
- moonlight through a window — cool key with hard-edged window pane shadow
- firelight from below — warm uplight, unsettling/ritual feel
- underwater caustics — rippling light patterns
Direction words
key from above, side-lit from the left, rim-lit from behind, uplit, top-down, across the frame.
Quality
soft, diffused, hard-edged, harsh, dappled, flickering, shafted, volumetric (with visible light beams / god rays).
Temperature & color
warm tungsten (~3200K), cool daylight (~5600K), icy blue moonlight, amber sunset, sodium-vapor street orange, cyan-magenta neon, pure white studio.
Camera, lens, angle, framing
Camera body / film stock
| Body / format | Effect |
|---|---|
| medium-format film | rich tonal depth, soft grain, editorial |
| 35mm film stock | classic cinematic grain |
| Fujifilm digital | accurate color science, slight warmth |
| Sony A7-series digital | clean modern look, high dynamic range |
| GoPro / action camera | wide distorted, immersive, fisheye-edge |
| disposable / single-use camera | flat flash, raw colors, 90s nostalgia |
| Polaroid / instant film | square frame, washed-out edges, faded |
| iPhone / smartphone snapshot | flatter, ubiquitous, casual realism |
| vintage Kodachrome | saturated reds, dense blacks |
| Lomography / cross-processed | shifted color casts, vignetting |
| security / CCTV camera | low resolution, timestamped, surveillance feel |
Lens / focal length
- 24mm wide-angle — landscape scale, slight edge distortion
- 35mm — natural reportage, documentary
- 50mm "nifty fifty" — human-eye perspective, classic portrait
- 85mm portrait lens — flattering compression, creamy background
- 135mm telephoto — strong compression, isolated subject
- macro lens — extreme close-up, every fiber visible
- fisheye — extreme spherical distortion, 180°
- tilt-shift — selective focus plane, miniature effect
- anamorphic lens — oval bokeh, horizontal flares, cinematic widescreen
Aperture / depth of field
- f/1.4 — extremely shallow — only one plane in focus, dreamy bokeh
- f/2.8 — shallow — subject crisp, background softly blurred
- f/5.6 — moderate — subject and immediate surroundings sharp
- f/8–11 — deep focus — most of the scene in focus, landscape standard
- f/22 — full depth — everything sharp foreground to infinity
Angle
- eye-level — neutral, journalistic
- low-angle — heroic, powerful, looking up
- high-angle / bird's-eye — diminished, vulnerable, surveying
- Dutch / canted angle — unease, tension
- over-the-shoulder — narrative POV
- worm's-eye view — extreme up
- top-down / flat lay — product, food, knolling
- aerial / drone shot — landscape scale
Framing / shot size
extreme close-up (eye only) → close-up (face) → medium close-up (head & shoulders) → medium shot (waist up) → cowboy shot (mid-thigh up) → full body → wide shot (subject in environment) → establishing / extreme wide shot (subject tiny).
Color & film stock
Color grades
- muted teal-and-orange — modern cinematic
- cool desaturated cyan-grey — Fincher / thriller
- warm sun-faded amber — nostalgic, summer memory
- bleach bypass — high-contrast, desaturated, gritty
- high-key clean white — fashion, clinical, optimistic
- low-key moody — dense blacks, single highlight
- technicolor saturated — vintage Hollywood, pop
- pastel washed-out — gentle, dreamlike
- monochrome / black-and-white with deep contrast
- duotone (specify two colors, e.g. deep navy and copper)
Film stock references
Kodak Portra 400 (creamy skin tones), Kodak Gold 200 (warm everyday), Fuji Velvia (saturated landscape), Ilford HP5 (B&W grain), CineStill 800T (tungsten halation, neon nights), Polaroid 600 (instant, faded).
Grain & texture
fine grain, coarse 35mm grain, digital noise, halation around highlights, light leaks, dust and scratches, vignette.
Materials & textures
Specifying material almost always upgrades a prompt. Replace generic nouns with material-noun pairs.
Fabric
navy wool tweed, crushed velvet burgundy, raw linen, iridescent silk, waxed canvas, aged denim, cashmere knit, neoprene, ripstop nylon, sheer organza.
Metal
brushed brass, patinated copper, polished chrome, gunmetal grey, raw cast iron, hammered pewter, anodized titanium, gold leaf, rusted steel.
Wood
raw oak, dark stained walnut, bleached pine, bamboo, driftwood, lacquered ebony, reclaimed barnwood, cherry wood with visible grain.
Stone & ceramic
Carrara marble with grey veining, terrazzo, raw concrete, smooth river stone, cracked terracotta, matte porcelain, glazed stoneware, quartz crystal.
Glass & plastic
frosted glass, clear acrylic, amber prescription glass, iridescent dichroic glass, crackled vintage glass, high-gloss lacquer, matte rubberized plastic.
Organic / skin / hair
freckled translucent skin, sun-weathered tan, visible skin texture and fine pores, wet glossy skin, coarse beard stubble, silver hair with individual strand definition, windblown loose hair.
Natural
moss-covered, frost-rimed, sand-blasted, salt-crusted, covered in fine dust, waterlogged, charred.
Mood / atmosphere
intimate, contemplative, foreboding, triumphant, melancholic, playful, uncanny, serene, kinetic, liminal, nostalgic, romantic, clinical, ritualistic, cozy, austere, ethereal, grounded.
Style families
Photography
editorial fashion, street photography, photojournalism, documentary, commercial product shot, food photography, architecture photography, landscape, wildlife, macro, long-exposure.
Illustration / paint
oil painting with thick impasto, watercolor wash, gouache, ink and brush, digital painting in the style of concept art, pen-and-ink crosshatch, children's-book illustration, vintage botanical illustration, Art Nouveau line art, Ukiyo-e woodblock print.
CG / 3D
high-fidelity 3D render, Pixar-style cel-shaded 3D, Blender Cycles render, Octane render, low-poly stylized, clay model, isometric 3D diorama.
Anime / comic
anime key visual, manga panel with screentone, Studio Ghibli–inspired painted backgrounds, American comic book ink and color, Moebius-style line art.
Other
vector illustration with flat colors, risograph print, blueprint / technical drawing, pixel art (16-bit), stop-motion claymation still, paper-cutout / collage.
Composition
- rule of thirds — subject on a third-line intersection
- centered symmetry — Wes Anderson balance
- leading lines — road, river, fence converging into the subject
- negative space — most of the frame empty around the subject
- layered depth — clear foreground, midground, background elements
- frame-within-a-frame — doorway, window, archway around subject
- over-the-shoulder POV, first-person POV, reflection (in mirror, water, glass)
- bokeh foreground / out-of-focus foreground element — depth cue
Words to avoid
These add token cost but no signal. Strip them from drafts:
- beautiful, gorgeous, stunning, amazing, breathtaking — meaningless on their own; if you mean a specific quality, name it (radiant skin, flawless symmetry, vivid saturation).
- high quality, best quality, masterpiece, 8k, ultra-detailed, hyper-realistic — these are SD/booru-era token-incantations; modern NL models do not need them and they often produce over-rendered plastic results. Use real descriptors of the medium instead (shot on 35mm film, fine micro-detail visible at close inspection).
- vibrant, dynamic, epic — replace with a concrete adjective.
- nice, good, cool, awesome — never appear in a professional prompt.
- trending on artstation, award-winning — legacy magic words, no longer help.
How to mix from this file
A strong draft typically lands like this — one item from each row:
| Row | Example pick |
|---|---|
| Style / medium | medium-format film photograph |
| Lighting setup | golden-hour backlight |
| Camera & lens | 85mm portrait lens at f/2.0 |
| Color grade | warm Kodak Portra palette |
| Material call-out | navy wool tweed, brushed brass |
| Mood | contemplative |
| Composition | medium close-up, rule-of-thirds |
If your draft is missing any row, that is the row to upgrade next.
Emotion conveyance (non-verbal)
Don't just name an emotion ("sad", "happy") — direct the body that produces it:
- Posture — shoulders rolled forward, weight on back foot, leaning toward / away from another subject, slumped against a wall, drawn upright, mid-step.
- Gaze — direct at camera, off-frame past the lens, downcast, soft-focus middle-distance, eyes closed, peripheral glance to another subject.
- Hand tension — relaxed open palm, white-knuckled grip, fingertips just touching a surface, fingers half-curled, hand half-raised mid-gesture.
- Micro-expression — slight asymmetric half-smile, tightened jaw, parted lips, raised inner brow, single tear-track without active crying.
- Inter-subject distance — close enough to share breath, an arm's length apart, separated by a single object, on opposite sides of the frame.
Environmental storytelling
Let the setting imply backstory without spelling it out:
- Wear patterns — scuffed floor in front of one chair, worn fingerprints on a single switch.
- Selective clutter — one half of the desk pristine, the other buried; a single chair pulled out from a long empty table.
- Aftermath cues — half-finished meal, overturned cup, recently extinguished candle, fresh footprints in dust.
- Anachronism — one modern object in a period scene, or vice versa.
Scale cues
Make scale legible by giving the model a reference:
- Tiny figure in vast space — a small silhouette against an oversized cathedral nave, dwarfed by a mountainside, ankle-deep in a sprawling tidal flat.
- Foreground reference object — a coffee cup, a hand, a coin, a doorway placed near the camera to anchor the size of the main subject.
- Forced perspective — exaggerated foreshortening with the subject filling the frame and the environment retreating.
Motion cues
Stills can imply motion through frozen physics:
- Long-exposure — light trails, smeared crowd, silken water.
- Frozen action — water droplets suspended mid-splash, fabric mid-swing, hair caught mid-toss, dust kicked up in an arc behind a running figure.
- Implied trajectory — leaning body weight, follow-through pose, displaced debris.
- Motion blur on subject vs sharp background (panning shot) or vice versa.
Illustration / design-specific controls
When the medium is illustration, comic, vector, or graphic design, reach for these instead of (or alongside) camera language:
- Line work — uniform clean linework, varied line weight emphasizing silhouette, sketchy hatching, no outlines (paint-only), tapered ink strokes.
- Shading technique — flat colors, cel-shaded with two tones (base + shadow), soft airbrushed gradient shading, cross-hatching, screentone dots, watercolor wash with bleed edges.
- Layout hierarchy — clear focal point, supporting elements arranged in a Z-pattern / circular flow, generous negative space framing the subject, grid-based layout for editorial / poster work.
- Negative space — leave the upper third clear for headline text, frame the subject with empty sky, use silhouette against flat ground.
- Palette discipline — limited two-color palette, analogous palette plus single accent, duotone, monochrome with one saturated highlight.
- Typography-safe regions (when the user does want text) — reserve a horizontal band along the bottom edge, a vertical column on the left, or a clean rectangle of low-contrast background.
- Vector / shape language — geometric flat shapes, rounded organic forms, sharp-cornered constructivist forms, isometric construction, pixel-grid-aligned for pixel art.
Example Prompts — Annotated
Reference these when calibrating tone, length, and structure. Each example shows the input intent (what the user asked for) and the delivered prompt, with brief notes on which formula and toolkit choices were applied.
Example 1 — Simple portrait, photographic
Intent: "A confident woman in a red coat for a fashion ad."
Prompt:
An editorial fashion photograph of a tall woman in her thirties wearing a tailored crimson wool overcoat, charcoal cigarette trousers, and pointed black ankle boots, standing with one hand in her pocket and a faint smile, slightly turned toward camera. Set against a seamless deep cherry-red studio backdrop. Medium-full shot, centered, 85mm lens at f/2.8, shallow depth of field. Three-point softbox lighting with a warm key from camera-right and a subtle rim from behind, shot on medium-format film, pronounced grain, saturated reds, cinematic warmth.
Notes:
- Formula §1 (text-to-image).
- Style/medium: editorial fashion photograph + medium-format film.
- Camera: 85mm f/2.8, medium-full, centered.
- Lighting: three-point softbox + warm key + rim.
- Material: tailored wool, pointed leather boots.
Example 2 — Cinematic narrative scene
Intent: "Lonely diner at night."
Prompt:
A long-exposure 35mm film photograph of an empty corner diner at 2 a.m., a single weary waitress in a faded teal uniform leaning on the counter and staring at her phone, an untouched coffee cup at her elbow. Wide medium shot through the rain-streaked plate-glass window from the sidewalk outside, 35mm lens at f/4. Cold blue ambient streetlight outside contrasting with the warm fluorescent buzz inside, soft halation around the interior lights, muted teal-and-amber color grade, fine 35mm grain, faint reflection of indistinct neon glow across the wet asphalt in the foreground.
Notes:
- Subject is singular (the waitress) with environmental story.
- Positive framing ("empty corner diner") rather than "no customers".
- Lighting names both sources (cold streetlight + warm fluorescent).
- Composition layered: foreground (wet asphalt + reflection), midground (window), subject behind.
Example 3 — Multi-reference product render
Intent: Two images provided — a wireframe sketch of a chair and a fabric swatch. User wants a finished render.
Prompt:
Using the attached wireframe sketch as the structural form and the attached oat-colored bouclé swatch as the upholstery texture, render a high-fidelity 3D product visualization of the resulting single-seater armchair. Place it in a sun-drenched minimalist Scandinavian living room with a pale oak floor and a single ceramic vase on a low side table. Three-quarter front view, 50mm lens, eye-level, shallow depth of field with the chair sharp and the background gently blurred. Soft warm key light from a large window camera-right, cool fill bouncing from a white wall on the left, neutral natural color palette with subtle warm highlights on the bouclé fibers.
Notes:
- Formula §2 (multimodal).
- Each reference's role stated explicitly: sketch → structure, swatch → texture.
- Setting and lighting added on top of the references.
Example 4 — Edit existing image
Intent: "Remove the bystander from my photo without messing up anything else."
Prompt:
Remove the man in the dark blue jacket standing on the right edge of the frame. Cleanly inpaint the space he occupied as a continuation of the brick wall and the cobblestones behind him. Keep the subject, her pose, her clothing, the lighting direction and quality, and all background elements unchanged from the source image; any existing signage on the left stays in place as it appears in the source but its text is not re-rendered or altered.
Notes:
- Formula §3 (editing).
- Names the target precisely ("man in the dark blue jacket on the right").
- Equal weight on what to PRESERVE.
- No new lighting / style invented; the edit is surgical.
Example 5 — Style transfer
Intent: "Make my street photo look like a Van Gogh."
Prompt:
Recreate the exact composition and content of the source photograph in the painted style of late Van Gogh: keep every building, the street layout, the figure positions, and the perspective unchanged, but render the scene with thick impasto brushwork, visible directional strokes following form, swirling textured sky, and a warm yellow / deep cobalt-blue palette with vivid contrast.
Notes:
- Formula §4 (style transfer).
- "Keep composition / content unchanged" is explicit so the model doesn't reinvent geometry.
- Style call-outs are specific: impasto, directional strokes, swirling sky, two-color palette.
Example 6 — Text inside the image (user explicitly asked)
Intent: "A poster that says 'NEW YORK' with the skyline showing through the letters."
Prompt:
A graphic typographic poster on a solid black background. The bold sans-serif letters "NEW YORK" fill the center of the frame in a single line, acting as a cut-out window: a photograph of the Manhattan skyline at golden hour is visible only inside the letterforms, with the rest of the canvas remaining pure black. Crisp, perfectly legible type. Tight centered composition, slight letterforms tracking, contemporary minimalist poster style.
Notes:
- Formula §6 (text-in-image).
- Exact words in quotes:
"NEW YORK". - Type style named: bold sans-serif.
- Effect described precisely (letters as cut-out window).
- Restriction stated ("rest of the canvas remaining pure black") to prevent the model from also painting a skyline outside the letters.
Example 7 — Character consistency set (3 shots)
Intent: "Three shots of the same detective character — at a desk, on a rooftop, in a car."
Canon block (reused):
A wiry man in his late forties with close-cropped salt-and-pepper hair, a three-day stubble, deep-set tired grey eyes, wearing a rumpled charcoal trench coat over a wrinkled white shirt with the top button undone and a loosened black tie, holding a worn leather notepad.
Shot 1 — desk:
[canon block]. Sitting alone at a cluttered wooden desk in a dim precinct office at night, leaning over a spread of crime-scene photographs, brow furrowed. Medium close-up, slightly high angle, 50mm lens at f/2.2. Single warm desk-lamp key from camera-left casting deep shadows across his face, cool blue light leaking in through window blinds in the background, muted teal-and-amber grade, 35mm grain.
Shot 2 — rooftop:
[canon block]. Standing at the edge of a rain-slick rooftop at dawn, looking out over a grey city skyline, breath visible in the cold air. Wide medium shot, slight low angle, 35mm lens at f/4. Cool overcast pre-sunrise key from camera-right, faint warm city lights below, desaturated blue-grey palette with a single warm note from a distant neon sign, fine grain.
Shot 3 — car:
[canon block]. Sitting behind the wheel of a stationary sedan parked on an empty avenue at night, half-lit by the dashboard glow, eyes scanning out the windshield. Close-up from the passenger seat, 35mm lens at f/1.8. Warm amber dashboard underlight on his face, cold cyan streetlight from outside the driver's window, strong chiaroscuro, deep blacks, anamorphic-flavored widescreen.
Notes:
- Formula §7 (consistency).
- Canon block is verbatim across all three; only setting/action/lighting vary.
- Each shot still picks lighting + camera + grade individually.
Example 8 — Tag-list input → NL rewrite
Intent (raw user input): "1girl, long blonde hair, blue eyes, white dress, beach, sunset, looking at viewer, smiling, cinematic, masterpiece"
Prompt (rewritten):
A warm cinematic photograph of a young woman with long flowing blonde hair and pale blue eyes, wearing a flowing white linen sundress that catches the breeze, standing barefoot on a stretch of damp sand at the water's edge, looking directly into the camera with a soft genuine smile. Wide medium shot, low angle close to the sand, 50mm lens at f/2.2, shallow depth of field. Golden-hour backlight from the setting sun behind her creating a halo through her hair and a long shadow stretching toward the camera, warm amber-and-rose palette, fine 35mm grain.
Notes:
- All tag-syntax residue removed.
- "1girl" → "young woman"; "looking at viewer" → "looking directly into the camera"; masterpiece / cinematic → replaced with real production language.
- Added specific lighting (golden-hour backlight) and one material call-out (linen sundress).
Example 9 — Vague input clarified
Intent (raw user input): "make me a cool image of a robot"
Skill response shape:
Quick question before I draft — what mood: clean futuristic / battle-worn / cute companion / unsettling uncanny? Want me to make a sensible default (battle-worn, cinematic) and you can redirect from there?
[On user OK or no response in single-turn mode, default and proceed:]
A cinematic medium close-up of a battle-worn humanoid robot standing in a fog-filled industrial yard at dusk, scratched matte-grey alloy plating streaked with rust and oil, exposed cabling running down one arm, a single intact glowing amber eye-lens fixed on something off-frame. Low angle, 35mm lens at f/2.0, shallow depth of field. Cool overcast key from above and a single warm sodium-vapor lamp from camera-right creating a strong rim along the wet metal, volumetric haze, desaturated steel-and-amber grade, fine grain.
Notes:
- Asked exactly one focused clarifying question before drafting.
- Stated the assumed default explicitly so the user can override.
- Final prompt still followed the full workflow.
Example 10 — Avoiding accidental text rendering
Intent: "A coffee shop interior" (user did NOT ask for text/signage).
Prompt:
An editorial-style photograph of a quiet independent coffee shop interior on a weekday morning, a single barista in a denim apron polishing a brass espresso machine behind a worn wooden counter, two empty stools in the foreground. Warm overhead pendant lights with exposed filament bulbs casting a soft golden glow across the room, daylight filtering in through a large front window out of frame to the left. Medium shot from across the counter, 35mm lens at f/2.8. Materials: reclaimed barnwood counter, brushed brass espresso machine, white subway-tile back wall, hanging dried-flower bouquets. Any signage, menu boards, or labels are softly out of focus and indistinct in the background. Warm earthy palette, fine 35mm grain.
Notes:
- Explicitly handled the text trap: "any signage… softly out of focus and indistinct" prevents the model from rendering hallucinated coffee-menu gibberish.
- Materials section reinforces tactile realism without resorting to "masterpiece quality" style filler.
Prompt Formulas by Scenario
Reach for the formula matching the user's task. All formulas produce flowing natural English, never tag lists or weight syntax.
1. Text-to-image (no reference images)
Formula:
[Style / medium] + [Subject + key descriptors] + [Action / expression]
+ [Setting / context] + [Composition / framing] + [Lighting] + [Color / texture]Template paragraph:
A[style/medium descriptor]of[subject with 2–3 vivid descriptors],[action or pose], set in[location with one or two environmental details].[Composition: framing + camera angle + lens/DOF].[Lighting: direction + quality + temperature], with[color palette and notable textures].
Worked example:
A cinematic medium-format film photograph of a weathered fisherman in a yellow oilskin coat, hauling a net over the gunwale of a small wooden boat, set against a churning slate-grey North Atlantic at dawn. Low-angle shot from the waterline with a 35mm lens at f/2.8, shallow depth of field. Cold, diffuse pre-sunrise light from the upper left, with cool steel-blue tones offset by the saturated yellow of the coat and the raw, salt-stained grain of the wood.
2. Multimodal generation (with reference images)
When the user attaches reference images.
Formula:
[Reference roles] + [Relationship instruction] + [New scenario] + [Style/composition direction]State explicitly which reference contributes what. Common roles: subject, pose, style, texture / material, color palette, background, layout/structure.
Template:
Using[reference A]as the[role of A]and[reference B]as the[role of B], render[new scenario]in[style/medium].[Optional: lighting + composition direction].
Worked example:
Using the attached napkin sketch as the structural layout and the attached linen swatch as the upholstery texture, render a high-fidelity studio product shot of a single-seater armchair in a sun-drenched minimalist living room. Three-point softbox lighting with warm key from camera-right, shot on a 50mm lens at f/4.
3. Image editing (conversational, no new references)
Treat the existing image as the unchanged baseline. The prompt should make crystal clear what changes and what must stay the same.
Formula:
[Edit operation: remove / add / replace / change] + [Target region or element]
+ [Desired result] + [What to preserve exactly]Worked examples:
Remove the man in the red jacket from the left side of the frame. Keep the lighting, the woman's pose and outfit, and the background unchanged from the source image; existing café signage stays in place but remains indistinct and is not re-rendered. Cleanly inpaint the space behind him as a continuation of the wet cobblestones.
Change the car's color from black to a deep burgundy with a subtle metallic flake. Preserve the exact lighting, reflections, and the rest of the scene unchanged.
4. Style transfer / composition (editing with new references)
Formula:
[Source: existing image] + [Reference: style image] + [What to keep from source]
+ [What to take from reference]Worked example:
Recreate the exact content and composition of the source photograph in the painterly style of the attached Van Gogh reference: keep the buildings, street layout, and figure positions exactly, but apply the reference's thick impasto brushwork, swirling sky, and warm-cool yellow-blue palette.
5. Real-time / web-search informed (when the runtime supports grounding)
When the model can pull live data.
Formula:
[Search / retrieval request] + [Analytical translation step]
+ [Visualization instruction]Worked example:
Search for the current weather and time in Reykjavík. Use that data to choose the lighting and atmospheric conditions of the scene. Visualize a miniature diorama of the city sitting inside a glass snow globe, rendered as a photorealistic studio product shot on a dark walnut desk, soft top-down key light.
6. Text rendered inside the image (only when user asks)
Default: do NOT add text. When the user explicitly wants legible text in the image:
Formula:
[Exact wording in double quotes] + [Typography style] + [Placement / hierarchy] + [Rest of scene]Rules:
- Quote the exact words:
the word "URBAN EXPLORER". - Name the type style: bold sans-serif, flowing brush script, condensed serif, hand-painted lettering, neon tube lettering. Naming a specific font (e.g. Century Gothic, Impact) helps stronger models.
- Place it deliberately: centered banner across the top, small caption in the lower-right corner, engraved into the metal plate.
- For multilingual rendering, name the target language explicitly: the word "歓迎" in elegant Japanese kanji brush script.
Worked example:
A high-end commercial beauty shot of a sleek nude-pink moisturizer jar on a warm beige studio background, soft diffused key light from above. Across the top, render the word "GLOW" in a flowing elegant brush-script font, centered. Below it, render "10% OFF" in a heavy blocky Impact-style font, and below that "Your First Order" in a thin minimalist Century Gothic. Keep all three lines crisp and legible.
7. Character or product consistency across a set
For storyboards, ad sets, or character sheets where the SAME subject must appear in multiple shots.
Formula:
[Character/product canon block] (reused verbatim across prompts)
+ [Per-shot: setting + action + framing + lighting]Write the canon block once — a tight 30–50 word description of fixed traits (hair, build, clothing signature, distinguishing marks, or product silhouette + material + color). Reuse it word-for-word in every prompt of the set, then vary only the scene-specific portion.
Worked example (canon block):
A tall lean woman in her late twenties with copper-red shoulder-length hair tucked behind one ear, pale freckled skin, sharp green eyes, wearing a navy oversized wool peacoat over a charcoal turtleneck and dark jeans, plain brown leather boots.
Then per-shot:
[canon block]. Sitting alone in a window seat of a near-empty Tokyo subway car at night, reading a worn paperback. Wide medium shot from across the aisle, 35mm lens, shallow depth of field. Cool fluorescent overhead light reflecting off the dark window beside her, muted teal and amber grade, fine 35mm grain.
Quick reference: scenario → formula
| User wants | Use formula |
|---|---|
| Make a prompt from scratch | §1 Text-to-image |
| Attach refs, generate new image | §2 Multimodal |
| Tweak an existing generated image | §3 Editing |
| Apply a style from another image | §4 Style transfer |
| Use live web data in the image | §5 Real-time |
| Put readable words in the image | §6 Text-in-image |
| Multiple shots of same character/product | §7 Consistency |