
Generate Image
- 42 installs
- 269 repo stars
- Updated June 11, 2026
- gupsammy/claudest
Generate Image is an agent skill that supplies mode-specific prompting patterns for photoreal, product, logo, illustration, and text-in-image outputs.
About
Generate Image is an agent skill that encodes capability-specific prompting patterns for image models during prompt crafting—load the relevant section in workflow step 2 when you need consistent visual output without trial-and-error chat. Solo builders use it while building landing pages, app store assets, brand kits, and marketing creatives where lens choice, lighting, isolation vs lifestyle product framing, and quoted text matter. It does not run generation itself; it teaches the agent how to phrase requests for photoreal scenes, e-commerce product isolation, cinematic hero frames, kawaii or poster styles, and advanced text-in-image layouts. Pair it with your image tool or API of choice in the same workflow. Intermediate complexity assumes you already pick a generator and iterate with follow-up edits rather than expecting one-shot perfection.
- Five capability modes: photorealistic scenes, product photography, logos & text, stylized illustration, text rendering
- Photoreal mode: lens, aperture, lighting direction, and mood spelled like a photographer brief
- Product modes: isolation (e-commerce), lifestyle, and hero shots with text-safe framing
- Logo and text: quoted strings, typography weight/style, placement, and iterative refinement
- Nano Banana–oriented text rendering tips (quoted copy, font characteristics, placement)
Generate Image by the numbers
- 42 all-time installs (skills.sh)
- Ranked #916 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/gupsammy/claudest --skill generate-imageAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 42 |
|---|---|
| repo stars | ★ 269 |
| Security audit | 2 / 3 scanners passed |
| Last updated | June 11, 2026 |
| Repository | gupsammy/claudest ↗ |
What it does
Craft mode-specific image-generation prompts for photoreal scenes, product shots, logos, illustrations, and legible text overlays.
Who is it for?
Best when you're producing storefront, social, or app visuals and want repeatable prompt recipes per visual mode.
Skip if: Skip if you only need SVG/code logos, automated batch pipelines, or image APIs wired without prompt craft.
When should I use this skill?
During prompt crafting workflow step 2 when loading mode-specific image prompting guidance.
What you get
You get structured, mode-aware prompts—camera, light, style, and quoted text—ready to paste into your image generator and refine.
- Mode-specific image prompts
- Iteration-ready text and typography briefs
By the numbers
- Five documented capability pattern sections (photorealistic, product, logos & text, stylized illustration, text renderin
Files
Requires GEMINI_API_KEY environment variable and uv package manager.
Workflow
1. Understand — Determine mode (t2i, i2i, multi-reference), gather parameters (model, aspect ratio, resolution, output path). If the prompt requires precise execution (specific pose, asymmetric framing, exact crop), default to --batch 3 or --batch 4 and surface this to the user — image generation is stochastic and precise directives hit ~50% per seed. Exit: mode, parameters, and batch size are clear. 2. Craft prompt — Default to the minimal prompt that can carry the intent: t2i uses narrative prose; i2i/multi-reference uses a reference block plus the minimal directive. Apply the Core checklist (always, for the matching mode). Reach into the Escalation toolkit only on a known-hard signature — a detail/geometry-fidelity shot — or after a batch shows drift; then add only the specific lock for the attribute that is drifting, not the whole kit. Over-constraining a simple edit degrades it as surely as under-specifying a complex one. Exit: prompt written, Core items satisfied, escalation tools added only where a signature or observed drift justifies them. 3. Confirm — Show the user the exact prompt, input images (if any), model, resolution, aspect ratio, and batch size. Ask for confirmation. Exit: user approves. 4. Generate — Run the script with confirmed parameters. Exit: images are saved and displayed. 5. Iterate — Present results and evaluate against intent before offering refinements. Evaluation order by mode: t2i — subject correctness, composition, style fidelity. i2i edit — the changed element looks right, nothing else changed. Multi-reference composition — the primary transferred attribute matches its source reference FIRST (for a detail shot, the construction geometry — width, edge shape, count, angle), secondary consistency (identity, environment) holds SECOND, staging (lighting, composition, framing) THIRD. Decide what's primary per task. Cherry-pick the winning frame from the batch rather than re-prompting for consistency past ~75%. Exit: user is satisfied or moves on.
Default Output & Logging
When the user doesn't specify a location, save images to:
~/Documents/generated images/Every generated image gets a companion .md file with the prompt and model used (e.g., logo.png → logo.md).
When gathering parameters (aspect ratio, resolution), offer the option to specify a custom output location.
---
Core Prompting Principle
Describe scenes narratively, not as keyword lists. Gemini's language model parses prose with full semantic understanding — narrative prompts encode spatial relationships, mood, and intent that comma-separated tags cannot express. Tag-style prompts lose compositional meaning and produce generic results.
Bad: "cat, wizard hat, magical, fantasy, 4k, detailed"
Good: "A fluffy orange tabby sits regally on a velvet cushion, wearing an ornate
purple wizard hat embroidered with silver stars. Soft candlelight illuminates
the scene from the left. The mood is whimsical yet dignified."Describe positively, never via negation. Every concept named in a prompt biases the output toward that concept — even when preceded by "not", "no", or "do not". Diffusion models condition on tokens regardless of polarity. To exclude X, either (a) name a positive alternative that fills the same role, or (b) scope the prompt so X has no place to land.
Bad: "A clean studio backdrop. No warm tones, no cream, no beige, no tan."
Good: "A clean cool-neutral gray studio backdrop with subtle blue undertones."
Bad: "A headshot with no harsh shadows on the face, no distracting background."
Good: "A headshot on a clean neutral gray backdrop, even soft frontal fill light
that flatters the face."This rule applies everywhere in the skill — t2i prompts, i2i directives, reference role descriptions, and framing instructions.
Name sources explicitly — leave no ambiguity in references. Every element in the prompt should trace to a specific source: "the man from Image 2" not "this man"; "the shirt from Image 3" not "the shirt". Ambiguous references bind to whichever source the model weights most, which is never reliably the right one. This isn't about over-describing — don't re-describe what the reference already shows. It's about making each reference point to exactly one source.
A useful formula: [Subject] doing [Action] in [Context]. [Camera/Composition]. [Lighting]. [Style]. [Constraint]. Not every prompt needs every element — match detail to intent. If the user has a specific vision, be prescriptive (exact descriptions); if exploring, be open (general direction, let the model decide details). Ask if unclear.
Advanced Prompting Techniques
Hyper-specificity: Be precise about quantities, positions, and attributes. "Three red apples arranged in a triangle on a wooden table" outperforms "some apples on a table." Every vague word is a degree of freedom the model fills arbitrarily.
Context and intent: State the purpose. "A hero image for a coffee brand landing page" produces different results than "a photo of coffee" even if the visual subject is the same, because intent shapes composition, mood, and framing.
Step-by-step instructions: For complex scenes, break the prompt into sequential directives. "Start with a wide desert landscape. Place a lone figure walking left-to-right in the lower third. Behind them, a massive sandstorm approaches from the right."
Exclusion via positive constraint: When something must be absent from the output, do not name it under a negation. Either name a positive alternative ("clean unbranded surface" instead of "no logos") or scope the scene so the unwanted element has no place to land ("a closed laptop on the desk" makes a screen impossible to render). Naming X under "no X" makes X more likely, not less.
Camera control: Specify shot type (extreme close-up, medium shot, aerial), lens (fisheye, telephoto), and camera angle (low angle, bird's eye, Dutch angle) to control framing precisely.
Editing with reference images follows different principles — see references/editing-guide.md.
---
Key Editing Principles
Editing prompts direct changes rather than describing scenes. Point to what the model can see; describe only what it cannot. Specify intentionally — every adjective, color word, or preservation clause beyond the minimum competes with the reference image and degrades fidelity. The reliable shape is a reference block plus one Replace directive — the verb's implicit scope handles preservation, no stop clause needed. Details in editing-guide.md.
For multi-reference work (3+ images), use per-reference role assignment: one sentence that assigns each reference its specific contribution ("the facade from Image 2; the car from Image 3; the sky and lighting from Image 1") — see editing-guide.md "Per-Reference Role Assignment".
Base image goes first in --input — it becomes Image 1 in the prompt. Gemini numbers images sequentially from input order. Reference block labels must match input order exactly.
Names invoke aesthetics directly — referencing "shot on Kodak Portra 400" produces its characteristic look more reliably than describing warm skin tones and pastel highlights.
---
References
Load the relevant reference during prompt crafting (workflow step 2):
- references/capability-patterns.md — mode-specific tips for photorealistic scenes, product photography, logos, stylized illustration, text rendering, and grounding
- references/editing-guide.md — edit grammar, reference blocks, directive structure, image ordering, semantic masking, character consistency
- references/style-reference.md — named aesthetics lexicon (film stocks, cameras, studios, artists, movements)
---
Configuration
Model Selection
| Nano Banana (default) | Nano Banana Pro | |
|---|---|---|
| Speed | Fast, high-volume | Slower, higher quality |
| Resolutions | 0.5K, 1K, 2K, 4K | 1K, 2K, 4K |
| Extra ratios | 1:4, 4:1, 1:8, 8:1 | — |
| Thinking mode | Yes (minimal/low/medium/high) | No |
| Image search grounding | Yes | No |
| Max references | 14 | 11 (6 objects + 5 characters) |
| Text rendering | Advanced | Standard |
Default to Nano Banana for most requests. Use Nano Banana Pro when the user explicitly asks for maximum quality or when Nano Banana results need refinement.
Aspect Ratios
Both models: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 Nano Banana only: 1:4, 4:1, 1:8, 8:1
Resolutions
- 0.5K (~512px) — fast preview (Nano Banana only)
- 1K (~1024px) — default, fast
- 2K (~2048px) — high quality
- 4K (~4096px) — maximum detail
Defaults: 1K resolution, batch 1, aspect ratio auto-detected from base image (first input, or 1:1 if no images). Use 0.5K for quick previews and iteration (Nano Banana only). Use 2K for higher quality requests, 4K only when high detail is explicitly needed.
Thinking Mode (Nano Banana only)
Nano Banana supports controllable thinking levels that improve complex prompt interpretation:
- minimal (default) — fastest, suitable for straightforward prompts
- low/medium — balanced reasoning for moderately complex scenes
- high — maximum reasoning for complex multi-element compositions, precise text rendering, or intricate spatial layouts
Use --thinking high when the prompt involves precise spatial relationships, multiple text elements, or detailed composition requirements. For i2i editing, thinking mode also helps with multi-reference composition (3+ images), precise text/sign placement on existing scenes, and complex spatial edits where element positioning matters.
---
Script Usage
One unified script handles all modes: t2i, i2i, and multi-reference composition. Nano Banana is the default model.
# Text-to-image (t2i) — uses Nano Banana by default
uv run ${CLAUDE_PLUGIN_ROOT}/skills/generate-image/scripts/generate.py --prompt "A serene mountain lake at dawn" --output landscape.png
# Nano Banana Pro model
uv run ${CLAUDE_PLUGIN_ROOT}/skills/generate-image/scripts/generate.py --prompt "A serene mountain lake at dawn" --output landscape.png --model pro
# Image-to-image editing (i2i)
uv run ${CLAUDE_PLUGIN_ROOT}/skills/generate-image/scripts/generate.py --prompt "Make it sunset colors" --input photo.png --output edited.png
# Multi-reference composition
uv run ${CLAUDE_PLUGIN_ROOT}/skills/generate-image/scripts/generate.py --prompt "Combine the cat from image 1 with the background from image 2" --input cat.png --input background.png --output composite.png
# With options (aspect ratio, resolution, thinking, batch, grounding, format)
uv run ${CLAUDE_PLUGIN_ROOT}/skills/generate-image/scripts/generate.py --prompt "Logo for 'Acme Corp'" --output logo.png --aspect 1:1 --resolution 2K --thinking highScript Options
| Flag | Short | Description |
|---|---|---|
--prompt | -p | Image description or edit instruction (required) |
--output | -o | Output file path (required) |
--input | -i | Input image(s) for editing/composition (repeatable, up to 14) |
--model | -m | Model: nano-banana (default) or pro |
--aspect | -a | Aspect ratio (auto-detects from base image / first input, or 1:1) |
--resolution | -r | Output resolution: 0.5K, 1K, 2K, or 4K (default: auto-detect or 1K) |
--grounding | -g | Enable Google Search web grounding |
--image-grounding | Enable image search grounding (Nano Banana only, use with --grounding) | |
--thinking | -t | Thinking level: minimal, low, medium, high (Nano Banana only) |
--quality | -q | Output compression quality 1-100 (JPEG only) |
--format | -f | Output format: png (default) or jpeg |
--batch | -b | Generate multiple variations: 1-4 (default: 1) |
--json | Output results as JSON for agent consumption | |
--quiet | Suppress progress output (MEDIA lines still printed) |
The script auto-detects resolution and aspect ratio from input images when flags are omitted, and automatically resizes large inputs (>2048px) before sending to the API.
---
Pre-Generation Checklist
Core items are the floor — apply them to every prompt of the matching mode. The Escalation toolkit is opt-in: skip it entirely for simple t2i and single-element edits. Reach in only on a known-hard signature (a detail/geometry-fidelity shot) or after a batch shows drift — and then add only the lock for the attribute that is actually drifting. Each added constraint costs fidelity on everything else, so escalation scales with how many independent things can drift, not with how ambitious the prompt is.
Core — t2i (always)
- [ ] Narrative description (not keyword list)?
- [ ] Positive framing throughout — no "no X" / "not X" / "do not X" clauses anywhere in the prompt?
- [ ] Camera/lighting details for photorealism?
- [ ] Text in quotes, font style described? (if the image has text)
- [ ] Aspect ratio appropriate for use case?
- [ ] Model choice appropriate? (Nano Banana default; Nano Banana Pro for max quality)
- [ ] Thinking level set for complex prompts? (Nano Banana only)
- [ ] Batch size matches precision needs? (
--batch 3or--batch 4for precise pose / framing / asymmetric directives)
Core — i2i / multi-reference (always)
- [ ] Reference block at start of prompt labeling each image's role?
- [ ] Reference roles positive-only — lists what to USE from each ref, never what to ignore? (see editing-guide.md "Reference Block")
- [ ] Minimal directive pattern? (Reference block + one Replace directive — no "do not change anything else" stop clause, no decorative preservation clauses)
- [ ] Positive framing throughout the directive? (no negation anywhere, including locks and composition clauses)
- [ ] Base image first in
--input(Image 1), and prompt labels match input order? (mislabeled roles cause character drift) - [ ] Only one change per prompt? (split competing directives into sequential passes)
- [ ] When extracting/transferring elements: explicitly named each element rather than generic "outfit/object from image X"?
- [ ] One reference carrying an element you need to preserve while another reference could compete with it? → add
CRITICAL — [element]proactively: assign the source reference explicitly ("Image 1 is the sole identity source. Image 2 is scoped to [attribute] only."). Do not wait for drift — role competition defeats text directives silently. - [ ] No color labels competing with reference image? (color words override visual reference — see editing-guide)
- [ ] Base image has minimal accessories that could contaminate? (bags, hats, sunglasses bleed into output)
- [ ] Reference count within model limits? (Nano Banana: 14, Nano Banana Pro: 11)
- [ ] For 3+ references: per-reference role-assignment sentence? (one sentence assigning each image its contribution — see editing-guide.md "Per-Reference Role Assignment")
- [ ] For photorealistic human shots: skin & finish locked? (the one near-universal CRITICAL section — generative skin drifts to plastic/retouched)
Escalation toolkit — reach for only on a hard signature or observed drift
- [ ] An attribute (identity, lighting/color, orientation, drape) drifting across the batch? Lock that specific attribute in its own
## CRITICAL —section, positively phrased — without over-constraining the stable ones. (see editing-guide.md "Constraint Locking with CRITICAL Sections") - [ ] Detail shot whose construction won't hold? Geometry lock via per-attribute enumeration, not generic "match exactly". (see capability-patterns.md "Geometry Lock for Detail Shots")
- [ ] Nano Banana +
--thinking highwith 3+ references and weak adherence? Add the inventory preamble. ("Silently inventory the design-critical details: ...") - [ ] Follow-up shot from the same set? Continuity assertion. ("from the same set as Image N: same subject, same setting, same light" — see editing-guide.md "Continuity Assertion")
- [ ] Campaign with a locked hero image? Collapse to two-input form — hero as bundle-source + the single new-attribute reference. (see editing-guide.md "Single-Reference Collapse")
Capability Patterns
Mode-specific prompting tips. Load the relevant section during prompt crafting (workflow step 2).
---
Photorealistic Scenes
Think like a photographer: describe lens, light, moment.
- Specify camera (85mm portrait, 24mm wide), aperture (f/1.8 bokeh, f/11 sharp throughout)
- Describe lighting direction and quality (golden hour from camera-left, three-point softbox)
- Include mood and format (serene, vertical portrait)
Product Photography
- Isolation: Clean white backdrop, soft even lighting, e-commerce ready
- Lifestyle: Product in use context, natural setting, aspirational but authentic
- Hero shots: Cinematic framing, dramatic lighting, space for text overlay
Logos & Text
- Put text in quotes:
'Morning Brew Coffee Co' - Describe typography: "clean bold sans-serif with generous letter-spacing"
- Specify color scheme, shape constraints, design intent
- Iterate with follow-up edits for refinement
Stylized Illustration
- Name the style: "kawaii-style sticker", "anime-influenced", "vintage travel poster"
- Describe design language: "bold outlines, flat colors, cel-shading"
- Include format constraints: "white background", "die-cut sticker format"
Text Rendering
Nano Banana has advanced text rendering capabilities. For best results:
- Put all text in single quotes within the prompt
- Describe font characteristics: weight, style, size relative to the image
- Specify text placement: "centered at the top," "bottom-right corner"
- For multiple text elements, describe each separately with position
- Use
--thinking highfor complex multi-line text or precise typography
Google Search Grounding
Enable with --grounding flag when real-time data helps (weather visualizations, current events infographics, real-world data charts).
Image search grounding (Nano Banana only): Add --image-grounding alongside --grounding to enable image search results as additional visual context. Useful when the model needs to reference real-world visuals (product designs, architectural styles, specific locations).
---
Best Practices
Hyper-Specificity
Vague prompts produce generic results. Every unspecified attribute becomes a random variable.
Vague: "A woman in a park"
Specific: "A 30-year-old woman with shoulder-length auburn hair sits cross-legged
on a green wool blanket in a sun-dappled oak grove, reading a hardcover
book. Late afternoon golden hour, shallow depth of field at f/2.0."Quantities, colors, materials, spatial positions, and named objects all reduce variance.
Context & Intent
State what the image is for. Purpose shapes composition, mood, and framing decisions.
Generic: "A flat white coffee on a marble counter"
With intent: "A hero image for an artisan coffee brand's homepage — a flat white
in a handmade ceramic cup on a marble counter, steam rising, soft
morning light from the left, negative space on the right for text overlay"Step-by-Step Instructions
Complex scenes benefit from sequential directives rather than a single compound sentence.
"Start with a wide establishing shot of a misty fjord at dawn.
In the foreground, place a wooden dock extending from the lower left.
A small red sailboat is moored at the dock's end.
Mountains fill the background, their peaks just catching the first golden light.
The water is perfectly still, creating mirror reflections."Positive Framing for Exclusions
Naming a concept under negation ("no X", "not X") biases the output toward X — diffusion models condition on tokens regardless of polarity. To exclude something, name a positive alternative that fills the same role, or scope the scene so the unwanted element is physically not there.
Bad: "A professional headshot on a neutral gray backdrop.
No distracting background elements, no visible logos or text,
no harsh shadows on the face."
Good: "A professional headshot on a clean seamless gray backdrop,
even soft frontal fill light that flatters the face, the
wall-to-floor falloff smooth and uncluttered."The Good version states what's there, not what isn't. "Clean seamless" implies absence of distraction. "Even soft frontal fill" implies absence of harsh shadows. The model never has to suppress a named concept.
Camera Control
Photographic terms give precise control over framing and perspective.
- Shot types: extreme close-up, close-up, medium shot, full shot, wide shot, extreme wide shot
- Angles: eye level, low angle (heroic), high angle (diminishing), bird's eye, worm's eye, Dutch angle
- Lenses: fisheye (distortion), wide-angle (expansive), normal 50mm (natural), telephoto (compression), macro (tiny subjects)
- Movement metaphors: "tracking shot following the subject," "slow dolly-in," "crane shot rising above"
---
Fashion & Garment Editing
Garment swaps and fashion compositing require specific techniques beyond generic i2i editing.
Base Image Selection
The base image matters as much as the prompt. Choose bases where:
- The garment being replaced is a contrasting color to the target (white base → olive swap, not olive → olive)
- The model/mannequin has minimal accessories (no bags, berets, sunglasses that bleed into output)
- The composition already has the target framing (Gemini cannot re-frame — see editing-guide.md)
Garment Swap Prompts
Use the reference block to label the image's role. Let the reference image carry color, cut, and texture — naming those attributes in the directive creates competing signals against the reference pixels.
Image 1: Base scene
Image 2: Reference shirt
Replace only the shirt on the mannequin with the blouse from Image 2.No stop clause — Replace already scopes the edit in place. See editing-guide.md "Minimal Directive Pattern".
Texture and Fabric
For premium fabric rendering, name the texture type without describing the color: "authentic linen texture with natural slub weave and organic drape." This gives the model rendering instructions while letting the reference image control color fidelity.
Multi-Step Fashion Edits
When changing outfit plus accessories or garment plus signage, split into passes (see editing-guide.md "Multi-Pass Editing"). Common two-step patterns:
- Garment swap first, then sign/easel text edit
- Outfit replacement first, then accessory adjustment
- Subject compositing first, then pose refinement
Multi-Reference Fashion Directive
The fashion instance of Per-Reference Role Assignment (editing-guide.md). When composing a full look from a face reference, separate garment references, and a lighting/backdrop plate, assign each reference its contribution in one positive sentence:
Perfectly replicate the exact features of the man's face from Image 2; the exact
shirt, trouser, belt, and shoe construction and color from Image 3; the cuff and
collar construction from Images 4 and 5; the cotton weave and sheen from Image 6;
and the lighting, backdrop, and photographic finish from Image 1.Each image gets a positively-named job, attributes are enumerated per image (not "the outfit"), and replication verbs ("perfectly replicate", "exact") carry positive intensity with zero negation. For the structural breakdown, see editing-guide.md "Per-Reference Role Assignment".
Detail Shots (fashion instance of Single-Reference Collapse)
This is the fashion application of Single-Reference Collapse (editing-guide.md). Once a hero image is locked for a campaign, every follow-up detail shot collapses to two-input form: the hero PNG as Image 1 (in this campaign the hero front shot was the bundle-source — it encoded the model's identity, the studio lighting, color science, backdrop, and wardrobe color), plus the single raw construction reference for the detail being shown (cuff macro, collar close-up, weave swatch). The character sheet and the separate lighting reference get dropped — the hero already carries what they contributed.
Inputs (order matters — base first):
Image 1: final/<colorway>/front.png ← bundle-source for everything locked this campaign
Image 2: raw/<colorway>/Cuff.jpg ← scoped to the construction detail being shown
Settings: --model nano-banana --thinking high --resolution 2K --batch 3 --aspect 3:4
Expected yield: 2/4 keepers (detail crops are simpler than full-body)(Adapt directory names to your campaign layout — `final/` and `raw/` above are MaisonX conventions.)
Which reference is the bundle-source is a per-campaign choice, not a fixed rule — here it was the hero front shot. This two-input form outperformed the original six-reference setup for detail shots because there were fewer competing signals. Anchor the detail shot to the bundle-source with a continuity assertion (editing-guide.md) and lock the construction with geometry enumeration (below).
Geometry Lock for Detail Shots
The fashion instance of Geometry Enumeration (editing-guide.md). Generic "match Image 2 exactly" does not preserve specific geometric attributes — the model interprets "exactly" aspirationally and width, edge shape, button count, and point spread all drift seed-to-seed. To lock garment geometry, enumerate each attribute explicitly in a dedicated CRITICAL — Geometry match section, named positively.
Canonical attributes by garment region:
| Region | Attributes to enumerate |
|---|---|
| Cuff | Width relative to wrist (e.g., 1.4–1.5x wrist circumference), edge shape (horizontal straight perpendicular to sleeve, sharp 90° corners), button count and placement (one at wrist edge, one on gauntlet placket above), topstitching gauge |
| Collar | Type (point / spread / cutaway / band / mandarin — name the one you want), point length (short / moderate / long), spread angle in degrees, stand height, top button position (visible at base of stand when fastened), placket type |
| Placket | Type (clean front / button-band / hidden), topstitching style (single-needle / double-needle), button count and spacing |
| Sleeve | Length (at the wrist / quarter-inch above / above the watch), drape (relaxed natural fold / pressed flat) |
Phrase the attributes in positive form. "Sharp 90° corners" not "not rounded". "Single button at the wrist edge" not "no double cuff". The rule from the Core Prompting Principle applies: every concept named under negation biases toward that concept.
Reference Orientation Lock
The fashion instance of scope completeness (editing-guide.md "Reference Block"). Scoping a reference to "construction only" can strip too much — Gemini drops the arm rotation, body angle, or camera direction that came baked into the reference's framing. The fix is to scope the reference to multiple positive attributes: construction AND orientation. Example for a back-of-wrist cuff shot:
Image 2 — Cuff construction and orientation reference. Silently inventory:
- The exact cuff geometry: [enumerated attributes above]
- The arm orientation: the model's torso is rotated so the back of the arm,
the back of the wrist, and the back of the hand face the camera. The cuff
button visible to camera sits on the outside of the wrist.Two positively-named contributions from one reference. The directive sentence that follows must echo both: "The cuff construction and the back-of-wrist orientation come from Image 2."
---
Working from Video References
When using reference videos as starting points for image generation (e.g., adapting an existing ad concept):
Frame Extraction
Use a two-pass approach with ffmpeg:
1. Scene detection — Find transition timestamps:
ffmpeg -i input.mp4 -vf "select='gt(scene,THRESHOLD)',showinfo" -vsync vfr -f null - 2>&1 | grep "pts_time"2. Targeted extraction — Extract a single frame at a specific timestamp:
ffmpeg -y -ss <TIMESTAMP> -i input.mp4 -frames:v 1 -update 1 output.pngStart with threshold 0.3 and lower to 0.15 if too few frames are detected. Fashion videos with smooth transitions (car wipes, camera pans) typically need the lower end.
Key Considerations
- Scene detection fires on visual composition changes, not semantic content changes. In videos where transitions are masked by passing objects, scene detection catches the transition itself, not the clean reveal after it. A second probe pass between detected timestamps is necessary.
- Always check
ffprobemetadata first (ffprobe -v quiet -print_format json -show_format -show_streams) to understand resolution, fps, and duration. - Name extracted frames descriptively (e.g.,
outfit_1_blue_denim.png) rather than by frame number — self-documenting folders save time during editing.
Editing & Composition Guide
Load this reference when the user provides input images for editing or multi-reference composition. These principles do not apply to text-to-image generation.
---
Core Principle
Connect the dots, don't describe. Reference images provide the visual context. The prompt's job is to tell the model what goes where — not re-describe what the model can already see. Over-describing creates competing signals between text and image that degrade output quality.
Bad: "Replace the front figure with a woman who has light fair skin, delicate oval
face, dark sunglasses, dark brown headscarf, sleeveless beige linen dress..."
Good: "Replace the front figure with the woman from image 3"The Pointing vs Naming Boundary
Point to reference images for elements that will be used whole (a face, a background, a scene). Explicitly name elements when extracting or transferring parts from a reference to a different context, because generic references like "outfit from image 2" don't tell the model which visual elements to isolate from surrounding context.
Bad: "Replace the clothing with the outfit from image 2"
Better: "Replace the clothing with the linen button-up shirt and wide-leg trousers from image 2"Name structural attributes (garment type, cut, fabric) when extracting parts from a reference — but never colors. Structural names tell the model which elements to isolate from the reference's surrounding context; colors belong to the reference image's pixels (see Color Labels Override Visual References below).
---
The Edit Grammar
Reference Block
Start multi-image edit prompts with a reference block that explicitly labels what each image represents. This disambiguates image roles before the directive.
Image 1: [base — canvas being edited]
Image 2: [reference role/description]
Image 3: [reference role/description]
[Main directive]Roles stay short — one phrase. Don't enumerate what the reference image contains in the label ("Image 2: Reference outfit — white linen blouse and camel trousers"). That re-describes what the pixels already show; the redundancy creates competing signals against the reference.
Scope references positively, never via exclusion. Phrases like "ignore the person, pose, face, and background" or "the model in this image is not retained" name the very elements you want the model to disregard, pulling them back into attention. Instead, list only what to use from each reference. The model treats anything you don't name as ambient context for the named contribution — no suppression required.
Bad: "Image 2: Reference for outfit. Ignore the person, pose, face, and background."
Good: "Image 2: Use only for the shirt — collar, button placket, chest pocket, fabric weave."
Bad: "Image 5: Collar reference. The face visible at the top is NOT the character."
Good: "Image 5: Use only for the collar construction — point shape, spread angle, stand height."The rule: list what to use; never list what to disregard. If a reference contains a contaminant (a face, a hat, a backdrop) the user doesn't want in the output, the way to suppress it is to specify the positive replacement elsewhere ("the identity comes from Image 1") and let the named-use scoping carry the rest.
*Scope to all the positive attributes you want — narrow scoping strips co-baked ones.* Scoping a reference to a single attribute ("use only for the silhouette") can drop other attributes baked into that reference's framing: the camera angle, the orientation, the lighting direction it happened to be shot at. If you need those too, name them — "use Image 2 for the silhouette and the 3/4 front camera angle". Two positive contributions from one reference, both explicit. See capability-patterns.md "Reference Orientation Lock" for a worked instance.
Minimal Directive Pattern
The whole pattern:
[Reference block]
Replace [scope] with [pointer to reference].One directive that points to the change. The verb's implicit scope carries the preservation work — Replace edits in place by definition, so everything else stays. No stop clause needed.
Example:
Image 1: Base photograph to edit
Image 2: Character and wardrobe reference
Replace the person in Image 1 with the woman from Image 2, wearing her exact outfit from Image 2.The directive's job is to connect the dots between images — name what's being swapped, point to where the replacement comes from. Add clauses beyond this only when the model is observably dropping a specific element you need to keep; in that case name only that single element in positive form ("Keep the seamless gray backdrop intact"), then stop.
Why minimal beats verbose. Every adjective, color word, or preservation enumeration is a degree of freedom the model reconciles against the reference image pixels. Text-vs-image conflicts get resolved by blending both signals — the result matches neither. Over-specifying what the reference already shows actively degrades fidelity. Add detail only when the model cannot infer it from the references and you have a specific outcome in mind for it.
No "do not change anything else" stop clause. Earlier versions of this skill used Do not change anything else. as the keystone of the minimal directive. Production evidence showed that clause activates the concept of changing other things — same negation-biases-toward-the-concept mechanism the Core Prompting Principle warns about. The verb's implicit scope is sufficient; the stop clause is gratuitous and counterproductive.
Verb choice.
- Replace anchors to the base scene (model edits in place). Use this for in-place edits.
- Change allows full recomposition — use only when you want the model to consider discarding the scene.
"Only" after the verb tightens scope: "replace only the front figure" beats "replace the front figure." Sentence order affects spatial placement — the element mentioned last in a spatial assignment tends to land in the more prominent position.
Per-Reference Role Assignment
For prompts with three or more references, a single positively-framed directive sentence outperforms a bulleted role list. The pattern: one sentence that assigns each reference its specific contribution, enumerating attributes positively, using replication verbs.
Canonical template:
Perfectly replicate the exact [element] from Image [N]; the exact [element]
from Image [M]; the [element] from Images [P and Q]; and the [environment,
lighting, and finish] from Image [S].Worked example (scene composite):
Perfectly replicate the exact building facade and signage from Image 2; the
parked vintage car from Image 3; the storefront awning from Image 4; and the
dusk sky, street lighting, and color grade from Image 1.What makes this work, structurally:
- Per-reference role assignment — every image gets a positively-named job, no ambiguity about which signal goes where
- Enumerated attributes per image — "facade + signage" not "the building"; name the specific parts, not the whole
- Replication verbs — "perfectly replicate", "exact" — positive intensifiers, zero negation
- Single coherent sentence — one directive carries the whole multi-ref assignment, easier for the model to parse than a bulleted block
When the prompt has 3+ references, default to this form. For a fashion worked example (face + garment construction + fabric + lighting across six references), see capability-patterns.md "Multi-Reference Fashion Directive".
Inventory Preamble (Nano Banana + thinking high)
For multi-reference composition with Nano Banana at --thinking high, prefix each reference's role with a "silently inventory" instruction. This forces Gemini's auto-regressive head to reason about each reference's design-critical details before diffusion activates — lifting adherence to the attributes you've assigned each reference measurably.
When to use:
- Multi-reference composition with 3+ references
- Nano Banana +
--thinking highonly (Pro has a fixed reasoning budget; no benefit) - Skip for single-reference edits (overkill) and t2i (irrelevant)
Prompt template (paste verbatim into your prompt):
## Reference inventory — analyze silently before generation
Image 1 — [Role: the bundle-source for X, Y, Z]. Silently inventory:
[the specific attributes this reference owns — enumerated positively].
Image 2 — [Role: the source for attribute W]. Silently inventory:
- [the precise geometry/structure of W — enumerated per attribute]
- [any co-baked attribute you also want from this reference, e.g. orientation]
[continue per reference]Follow the inventory block with a single positive directive sentence (Per-Reference Role Assignment, above) and any CRITICAL sections needed for geometry lock or continuity. For a worked fashion inventory (a reference scoped to both cuff construction and arm orientation), see capability-patterns.md "Reference Orientation Lock".
Constraint Locking with CRITICAL Sections
A ## CRITICAL — [attribute] block is a general constraint-locking primitive, not a fixed menu. Any attribute the model tends to drift on can get its own CRITICAL section that names the target positively and in detail. Geometry and continuity (below) are the two with full worked treatment, but production prompts routinely stack several different locks in one prompt:
| Lock | What it pins | One-line shape |
|---|---|---|
| Identity | a character's face/hair/skin across shots | "The subject is exactly the person from Image 1 — same hair, jawline, skin tone and grain; identity sourced exclusively from Image 1." |
| Skin & finish | texture realism, no plastic/waxy look | "Natural unedited skin with visible pores and matte complexion. Muted naturalistic neutral palette." |
| Lighting & color | photographic register continuity | "Same key-light direction, shadow density and falloff, color science, and finish as Image 1." |
| Geometry | exact construction / proportions | enumerate per attribute — see Geometry Enumeration |
| Orientation / pose | which way the subject faces | "The back of the head faces the lens; the head turns slightly so a sliver of cheek shows in profile at the frame's right edge." |
| Continuity | a whole bundle from one reference | "from the same set as Image 1: same subject, same setting, same light" — see Continuity Assertion |
| Subject-pose adaptation | drape/fall follows the new pose, not the reference's | "Adapt the garment drape to the described pose, not Image 3's pose. Follow Image 3 only for color, cut, construction." |
It's a balance. Each CRITICAL section removes a degree of freedom — raising fidelity on that attribute, but also adding tokens the model must reconcile and risking crowding out others. Lock the attributes that both (a) matter for this shot and (b) the model is actually drifting on. A simple t2i or single-element edit needs none. A complex multi-reference composition — where identity, construction, lighting, and skin realism all have to hold at once — earns several stacked locks. The count scales with how many independent things can drift, not with how ambitious the prompt is. Add a lock when you observe drift; don't pre-emptively lock everything. An attribute the references already carry reliably (because a hero image encodes it) needs no lock — see Single-Reference Collapse.
Two empirical patterns from production:
- Skin & finish is the near-universal lock for photorealistic human shots. Generative skin reliably drifts toward plastic, waxy, or over-retouched, so it earns a lock even when every other attribute is stable. It is often the only lock a hero-anchored shot needs.
- The harder it is for the references to carry an attribute, the more it needs an explicit lock. A back view where the face is mostly hidden has almost no pixels to anchor identity, so identity must be re-asserted in text even though a hero portrait exists. When a reference can't be used for a lock (e.g. a construction reference that leaks an unwanted face), describe the attribute in text instead and annotate the header —
## CRITICAL — Collar geometry (described, not image-referenced).
Place CRITICAL sections after the generation directive and shot description, one per attribute, each headed ## CRITICAL — [attribute]. Phrase every lock positively (the rule from the Core Prompting Principle in SKILL.md).
Continuity Assertion (Composition Lock)
When an output should inherit a bundle of qualities from one reference — an identity, OR a product, OR an environment plus its light, grade, and finish — assert sameness in a single clause instead of enumerating each quality. The bundle-source can be any reference; name the attributes that reference owns.
This is a [shot] from the same [session / set / scene] as Image [N]:
same [subject], same [setting], same [light].The phrase does enormous load-bearing work — production runs replaced multi-paragraph preservation blocks with this single sentence and got tighter coherence. Use it as a dedicated ## CRITICAL — Continuity section in any follow-up shot that should read as part of the same set as an earlier image.
Generic example — compositing a product onto a recurring set:
## CRITICAL — Continuity
This is a detail shot from the same set as Image 1: same product, same marble
table, same softbox lighting, same color grade. The bottle is the visual hero,
centered in the frame against the table surface.This works because asserting continuity ("same set as Image 1") re-anchors the whole bundle in one token-cheap clause, where re-describing each quality separately would multiply the degrees of freedom the model has to reconcile against the reference. For a fashion worked example (back-of-wrist detail anchored to a hero shoot), see capability-patterns.md "Detail Shots".
Single-Reference Collapse
Once one image reliably encodes a locked bundle of attributes — because an earlier generation produced it, or because it was assembled for the purpose — drop the references that originally contributed those attributes. A composite built from six references to establish a subject can collapse to two inputs for follow-ups: the locked image as the bundle-source, plus the single reference for the new attribute being introduced.
Fewer references means fewer competing signals, which raises fidelity on the attributes that matter. The collapse is not domain-specific:
- Character work — once a hero portrait locks the face, drop the turnaround/character sheets; use the portrait plus the new pose or wardrobe reference.
- Product work — once a hero render locks materials and lighting, drop the CAD/lighting refs; use the render plus the new angle reference.
- Scene work — once an establishing frame locks the environment and grade, drop the mood-board refs; use the frame plus the new foreground element.
Which reference becomes the bundle-source is a per-task choice, not a fixed rule. For the fashion instance (a hero front shot collapses the character sheet and lighting reference for every follow-up detail crop), see capability-patterns.md "Detail Shots".
Geometry Enumeration
A generic "match Image 2 exactly" does not preserve specific geometric attributes — the model reads "exactly" aspirationally and lets dimensions, edge shapes, counts, and angles drift seed-to-seed. To lock geometry, enumerate each attribute explicitly and positively in a dedicated ## CRITICAL — Geometry match section.
This applies to any object with specific geometry the reference cannot be trusted to carry on its own:
- A logo — stroke weight, corner radius, letter-spacing, x-height, overall aspect ratio
- An architectural feature — window proportions, column count, roof pitch, bay spacing
- A product silhouette — neck-to-body ratio, shoulder curve, base diameter
- A face — feature spacing, jaw angle, brow position
Phrase every attribute positively — "sharp 90° corners" not "not rounded", "a single fastener at the edge" not "no double row". For the fashion garment-region instance (cuff / collar / placket / sleeve attribute table), see capability-patterns.md "Geometry Lock for Detail Shots".
Image Ordering & Numbering
Base image (the canvas being edited) goes first in the --input list — it becomes Image 1 in the prompt, the most natural labeling. Reference images follow in order:
--input base.jpg -> Image 1 in prompt
--input ref_a.jpg -> Image 2 in prompt
--input ref_b.jpg -> Image 3 in promptAspect-ratio auto-detection reads from the first input (the base), so the output matches your canvas without needing --aspect. Label your reference block to match input order exactly — mismatched labels cause role confusion and character drift even when the directive is otherwise correct.
---
Multi-Image Composition
Reference images provide visual context — the prompt connects them. Point to images by number, assign elements to positions, and describe only what the model cannot infer from the images themselves. Nano Banana supports up to 14 reference images; Nano Banana Pro supports up to 11 (6 objects + 5 character references).
---
Character Consistency
- Use follow-up edits for multiple views of the same character
- Reference distinctive features explicitly in follow-ups
- Include "exact same character" or "maintain all design details"
- Save successful designs as reference for future prompts
---
Semantic Masking
No manual masking needed. Language creates the edit boundary — name the element to define the mask, specify the replacement, constrain the scope:
"Using the provided image of a living room, change only the blue sofa
to a vintage brown leather chesterfield.""Only" after the verb defines scope. The element name ("the blue sofa") defines the mask region. The change verb's implicit scope protects everything else — no stop clause needed.
---
Editing Failure Modes
Common ways i2i edits fail and how to avoid them. These patterns emerged from real production campaigns and apply to any editing workflow.
Color Labels Override Visual References
Every color word in an editing prompt creates a competing signal against the reference image. If you name a color ("rust", "terracotta", "olive green"), the model generates the text-defined tone rather than the actual shade visible in the reference. This happens because the model resolves text-vs-image conflicts by blending both signals — the result matches neither.
Remove all color words from editing directives. The reference block labels the image's role ("Reference shirt — women's linen blouse"); the directive says what to change ("Replace only the shirt with the blouse from the reference"). The reference image itself is the color spec.
This applies to any visual attribute already present in the reference: color, cut, texture, proportion. Naming these in text creates drift. If you want something replicated exactly, don't describe it — let the image be the sole authority.
High-Contrast Swap Targets
When replacing an element with something of a similar color (olive shirt → olive shirt of different cut), the model can't distinguish source from target and produces near-identical output. The fix is to choose a base image where the element being replaced is a contrasting color — e.g., use a white shirt base for an olive shirt swap. The high contrast gives the model an unambiguous replacement target.
Gemini Cannot Re-Frame
Gemini cannot execute virtual camera moves. Prompts like "zoom in on the storefront," "show this from a closer angle," or "crop to a tighter shot" will either reproduce the original composition or generate an inconsistent scene — they will not produce a re-framed version of the same content.
Always edit on a base image that already has the target angle, framing, and composition. If you need a closer shot, find or extract a frame at that angle rather than trying to prompt a re-frame.
Multi-Pass Editing
The minimal directive pattern (one Replace, no stop clause — see "Minimal Directive Pattern" above) works for a single change. When a prompt has two competing changes — garment swap plus sign text, outfit plus accessories, subject plus background — the model compromises on one.
Split into sequential passes: one change per generation call. Pattern for element replacement with correct proportions:
1. Pass 1: Remove the element entirely (e.g., "Remove the price sign and easel from the scene") 2. Pass 2: Re-add it using a visual reference (e.g., "Add the exact price sign and easel from image 2 to the right side of the shelf")
This remove-then-re-add approach is specifically important for text and sign elements, where text-only swaps change the words but distort the element's proportions and positioning.
Base Scene Contamination
Accessories and distinctive elements in the base image bleed into garment-swap outputs. If the base scene has a beret, the generated outfit may include a beret. If the base scene has sunglasses, they appear on the output model. Similarly, outfit references shot on plain studio backgrounds can override the base scene's location background — the gray studio backdrop replaces the street.
When choosing a base image for garment editing:
- Prefer images with minimal distinctive accessories
- Avoid bases where the model wears items (bags, hats, jewelry) you don't want in the output
- If using outfit references from studio shoots, verify the output preserves the base scene's environment
Pose to Make Unwanted Elements Impossible
When framing or composition language alone fails to suppress an unwanted element (positive crop instructions get interpreted aspirationally — "tight crop to forearm" still includes the belt), pose the subject so the unwanted element is physically not in the shot. The model gets a concrete physical solution instead of a suppression task.
Worked example: a cuff macro that kept including the trouser. The fix was not "no trouser in frame" — it was repositioning the arm:
Bad: "Crop tightly to the forearm and cuff only. No belt, no trouser visible."
Good: "The arm is held slightly away from the body so the cuff is silhouetted
against the studio backdrop. The forearm and hand fill the frame against
the seamless backdrop behind them."The trouser is suppressed not by naming its absence but by giving the arm a physical position where the trouser cannot be visible. Same technique works for accessories ("the hands hang at the sides, fingers loose" makes hand-in-pocket impossible), facial features ("the head is turned profile to camera" makes both eyes impossible to show), and backdrops.
Asymmetric Directive Collapse
Gemini's editorial training prior couples contrapposto with symmetric framing — "one hand in pocket, the other at side" gets symmetrized to "both hands in pockets" on ~50% of seeds. Standard hedging language ("naturally asymmetric") doesn't help; the model defaults to its prior.
To improve compliance, name the asymmetry as the explicit goal and describe both sides positively in separate clauses:
Bad: "Hands settle naturally — one in pocket, one at the side."
Good: "Only the left hand is in the pocket. The right hand rests visibly at the
right thigh with fingers loose. The arms are deliberately asymmetric."Naming both sides removes the option of guessing; describing the asymmetry as the goal removes the option of regression toward symmetry.
Listed Alternatives Collapse
When a prompt offers the model multiple options for the same attribute ("hands at sides, OR hand in pocket, OR adjusting cuff"), Gemini picks one option and repeats it across every seed in the batch — defeating the purpose of offering alternatives for variety.
For pose or styling variety across a batch, run separate generations each with a single distinct directive. One prompt = one option; the batch flag controls per-prompt variance, not directive variance.
Per-Seed Variance
Image generation is stochastic — the same prompt with the same references produces meaningfully different outputs across seeds. For directives requiring precise execution (specific pose, asymmetric framing, exact crop, geometry lock), expect ~50% per-seed hit rate even with locked recipes. The remedy is --batch 3 or --batch 4 plus cherry-picking the obeying frame, not iterating prompt language to push consistency past ~75%.
When cherry-picking, judge against the primary transferred attribute first — whatever the shot exists to deliver — then secondary consistency (identity, environment), then staging (composition, framing). Decide what is primary per task before you generate; for a multi-reference detail shot the primary attribute is usually the construction being highlighted, not the identity carrying it.
Style Reference Lexicon
On-demand reference for aesthetic naming. Use these names in prompts to invoke specific visual characteristics.
---
Film Stocks
| Film | Effect |
|---|---|
| Kodak Portra 160/400/800 | Warm skin tones, pastel highlights, fine grain, nostalgic |
| Kodak Ektar 100 | Vivid colors, high saturation, fine grain |
| Kodak Gold 200 | Consumer warmth, golden tones, everyday nostalgia |
| Fujifilm Pro 400H | Soft pastels, muted tones, airy highlights |
| Fujifilm Velvia 50 | Saturated colors, high contrast, punchy greens/blues |
| Fujifilm Superia | Consumer look, versatile, slight green cast |
| Cinestill 50D/800T | Cinematic, halation around lights, film-set feel |
| Ilford HP5 | B&W, medium contrast, versatile grain |
| Ilford Delta 3200 | B&W, high ISO, pronounced grain, night scenes |
| Kodak Tri-X 400 | B&W, classic photojournalism, gritty texture |
| Kodak Vision3 500T | Tungsten cinema film, blue shadows, warm highlights |
Modifiers: "pushed to 1600" (more grain, contrast), "cross-processed" (color shifts)
---
Camera Systems
| Camera | Association |
|---|---|
| Hasselblad 500C/X1D/X2D | Medium format, exceptional detail, fashion/portrait |
| Leica M6/M10/Q2 | Street photography, micro-contrast, documentary feel |
| Phase One IQ4 | Ultra-high resolution, commercial, studio |
| Mamiya RZ67 | Portrait, medium format, classic rendering |
| Contax T2/G2 | 90s aesthetic, Zeiss rendering, compact luxury |
| Fujifilm X100V | Modern street, Fuji colors, retro design |
| Canon AE-1 | 70s-80s consumer photography |
| Nikon FM2 | Mechanical reliability, photojournalism |
| Rolleiflex 2.8F | Twin-lens, square format, vintage portrait |
| Pentax 67 | "Medium format look," portrait favorite |
---
Lenses
| Lens/Type | Effect |
|---|---|
| Zeiss Otus 85mm | Razor sharp, smooth bokeh, clinical precision |
| Helios 44-2 | Swirly bokeh, vintage Soviet character |
| Leica Summilux | Soft glow wide open, classic rendering |
| Anamorphic | Oval bokeh, horizontal flares, cinematic 2.39:1 feel |
| Tilt-shift | Miniature effect, selective focus plane |
---
Animation & Studios
| Reference | Style |
|---|---|
| Studio Ghibli | Lush nature, soft lighting, whimsical, hand-painted feel |
| Makoto Shinkai | Hyper-detailed skies, light rays, emotional realism |
| Spider-Verse | Ben-Day dots, chromatic aberration, comic frames |
| Arcane (Netflix) | Painterly textures, dramatic lighting, stylized realism |
| MAPPA/Ufotable | Modern anime, dynamic action, high production value |
| Trigger | Exaggerated motion, bold colors, kinetic energy |
| Disney Renaissance | Clean lines, expressive characters, rich colors |
| Pixar | 3D polish, emotional storytelling, detailed textures |
| Satoshi Kon | Psychological, reality-bending, adult themes |
| Fleischer Studios | 1930s rubber hose, surreal, black and white |
| Cartoon Network/Adult Swim | Varied, often irreverent, stylized |
---
Illustration & Comics
| Artist/Style | Effect |
|---|---|
| Moebius / Jean Giraud | Intricate linework, surreal sci-fi landscapes |
| Alphonse Mucha | Art Nouveau, ornate borders, flowing forms |
| Frank Frazetta | Fantasy, muscular figures, dramatic action |
| Yoshitaka Amano | Ethereal, delicate lines, Final Fantasy aesthetic |
| Kim Jung Gi | Dense, ink-heavy, no preliminary sketches |
| Mike Mignola | Heavy shadows, angular, Hellboy style |
| Ukiyo-e | Japanese woodblock, flat colors, bold outlines |
| N.C. Wyeth | Golden age illustration, adventure, painterly |
---
Game Art
| Reference | Style |
|---|---|
| Breath of the Wild | Soft cel-shading, painterly skies, open world |
| Dark Souls/FromSoftware | Dark fantasy, intricate armor, oppressive mood |
| Hollow Knight | Hand-drawn 2D, bug aesthetic, moody underground |
| Cuphead | 1930s Fleischer, hand-drawn animation, jazz age |
| Cyberpunk 2077 | Neon-soaked, high-tech dystopia, chrome |
| Vanillaware | Painterly 2D, rich colors, fantasy (Odin Sphere) |
| 16-bit SNES/Genesis | Limited palette, chunky pixels, nostalgic |
| MS-DOS aesthetic | CGA/EGA colors, dithering, early PC |
---
Graphic Design & Eras
| Movement/Era | Characteristics |
|---|---|
| Swiss International Style | Grid-based, clean sans-serif, asymmetric |
| Bauhaus | Geometric shapes, primary colors, functional |
| Art Deco | Geometric elegance, gold/black, 1920s glamour |
| Psychedelic 60s | Warped letterforms, vibrant, Art Nouveau influence |
| Y2K aesthetic | Chrome, gradients, futuristic optimism |
| Vaporwave | 80s nostalgia, pink/cyan, Roman busts, glitch |
| Brutalist | Raw, unconventional layouts, stark typography |
---
Fine Art Movements
| Movement | Visual Language |
|---|---|
| Impressionism | Visible brushstrokes, light/color focus, en plein air |
| Post-Impressionism | Van Gogh swirls, Cézanne geometry, expressive color |
| Art Nouveau | Organic curves, nature motifs, decorative |
| Surrealism | Dreamlike, impossible scenes, subconscious imagery |
| Pop Art | Bold colors, commercial imagery, Warhol/Lichtenstein |
| Baroque/Caravaggio | Dramatic chiaroscuro, theatrical lighting |
| Pre-Raphaelite | Vivid colors, medieval subjects, intricate detail |
| Hudson River School | Romantic American wilderness, dramatic light |
| De Stijl/Mondrian | Primary colors, black grid, geometric abstraction |
---
Photographers & Directors
| Name | Known For |
|---|---|
| Gregory Crewdson | Cinematic suburban tableaux, uncanny |
| Annie Leibovitz | Editorial portraits, dramatic, intimate |
| Wes Anderson | Symmetry, pastels, whimsy, centered compositions |
| Roger Deakins | Naturalistic cinematography, atmospheric |
| Terrence Malick | Magic hour, existential beauty, nature |
| Junji Ito | Horror manga, spiral patterns, body horror |
| Hiroshi Nagai | Japanese city pop, nostalgic summer, poolside |
| Zdzisław Beksiński | Surreal, nightmarish, organic decay |
| William Eggleston | Saturated color, mundane subjects elevated |
| Saul Leiter | Abstract color street photography, painterly |
---
Quick Combinations
These patterns work well together:
Portraits: "shot on Kodak Portra 400, Hasselblad medium format, soft window light"
Street: "Leica M10, 35mm Summicron, Cinestill 800T at night"
Product: "Phase One IQ4, studio softbox lighting, commercial quality"
Fantasy: "Studio Ghibli style, Makoto Shinkai lighting"
Horror: "Junji Ito aesthetic, high contrast black and white ink"
Retro: "Y2K aesthetic, chrome text, gradient background"The lexicon is vast—these are samples. If you know a specific name, try it directly.
#!/usr/bin/env python3
# /// script
# requires-python = ">=3.10"
# dependencies = [
# "google-genai>=1.0.0",
# "pillow>=10.0.0",
# ]
# ///
"""
Generate and edit images using Nano Banana (Gemini 3.1 Flash Image) and Nano Banana Pro (Gemini 3 Pro Image).
Supports three modes:
- t2i (text-to-image): Generate from prompt only
- i2i (image-to-image): Edit a single image with a prompt
- Multi-reference: Compose from multiple images
Usage:
# Text-to-image (Nano Banana default)
uv run generate.py --prompt "A cat in space" --output cat.png
# Use Nano Banana Pro model
uv run generate.py --prompt "A cat in space" --output cat.png --model pro
# Image editing
uv run generate.py --prompt "Make it blue" --input photo.png --output edited.png
# Multi-reference composition
uv run generate.py --prompt "Combine cat from first with background from second" \\
--input cat.png --input background.png --output composite.png
# Batch generation (up to 4 images, async parallel)
uv run generate.py --prompt "A cat in space" --output cat.png --batch 4
# JSON output for agent consumption
uv run generate.py --prompt "A cat in space" --output cat.png --json
Options:
--prompt, -p Image description or edit instruction (required)
--output, -o Output file path (required)
--input, -i Input image path for editing (can be repeated up to 14 times)
--model, -m Model: nano-banana (default) or pro
--aspect, -a Aspect ratio (1:1, 16:9, 9:16, etc.)
--resolution, -r Resolution: 0.5K, 1K, 2K, 4K (default: auto-detect or 1K)
--grounding, -g Enable Google Search web grounding
--image-grounding Enable image search grounding (Nano Banana only)
--thinking, -t Thinking level: minimal, low, medium, high (Nano Banana only)
--quality, -q Output compression quality 1-100 (JPEG only)
--format, -f Output format: png (default) or jpeg
--batch, -b Generate multiple variations (1-4, default: 1)
--json Output results as JSON for agent consumption
--quiet Suppress progress output (MEDIA lines still printed)
Environment:
GEMINI_API_KEY - Required API key
Exit codes:
0 - Success
1 - Generation or validation error
2 - Environment error (missing API key)
"""
from __future__ import annotations
import argparse
import asyncio
import json
import os
import sys
from datetime import datetime
from io import BytesIO
from pathlib import Path
# --- Model registry ---
MODELS = {
"nano-banana": "gemini-3.1-flash-image",
"pro": "gemini-3-pro-image",
}
MODEL_DISPLAY = {
"nano-banana": "Nano Banana",
"pro": "Nano Banana Pro",
}
DEFAULT_MODEL = "nano-banana"
PRO_RATIOS = ["1:1", "2:3", "3:2", "3:4", "4:3", "4:5", "5:4", "9:16", "16:9", "21:9"]
NB_RATIOS = PRO_RATIOS + ["1:4", "4:1", "1:8", "8:1"]
PRO_RESOLUTIONS = ["1K", "2K", "4K"]
NB_RESOLUTIONS = ["0.5K"] + PRO_RESOLUTIONS
# --- Aspect ratio auto-detection ---
SUPPORTED_RATIOS = [
("1:1", 1.0),
("1:4", 1/4),
("1:8", 1/8),
("2:3", 2/3),
("3:2", 3/2),
("3:4", 3/4),
("4:1", 4/1),
("4:3", 4/3),
("4:5", 4/5),
("5:4", 5/4),
("8:1", 8/1),
("9:16", 9/16),
("16:9", 16/9),
("21:9", 21/9),
]
def get_closest_aspect_ratio(width: int, height: int, model: str = "nano-banana") -> str:
"""Find closest supported aspect ratio for given dimensions and model."""
valid_ratios = NB_RATIOS if model == "nano-banana" else PRO_RATIOS
candidates = [(name, val) for name, val in SUPPORTED_RATIOS if name in valid_ratios]
actual_ratio = width / height
closest = min(candidates, key=lambda x: abs(x[1] - actual_ratio))
return closest[0]
# --- Resolution auto-detection ---
def detect_resolution(images: list) -> str:
"""Auto-detect appropriate resolution from input image dimensions."""
if not images:
return "1K"
max_dim = 0
for img in images:
width, height = img.size
max_dim = max(max_dim, width, height)
if max_dim >= 3000:
return "4K"
elif max_dim >= 1500:
return "2K"
return "1K"
# --- Validation ---
def validate_model_params(model: str, aspect: str | None, resolution: str | None,
thinking: str | None, image_grounding: bool) -> None:
"""Validate parameters against model capabilities. Exits on invalid combinations."""
valid_ratios = NB_RATIOS if model == "nano-banana" else PRO_RATIOS
valid_resolutions = NB_RESOLUTIONS if model == "nano-banana" else PRO_RESOLUTIONS
if aspect and aspect not in valid_ratios:
print(f"Error: Aspect ratio '{aspect}' not supported by {model} model. "
f"Valid: {', '.join(valid_ratios)}", file=sys.stderr)
sys.exit(1)
if resolution and resolution not in valid_resolutions:
print(f"Error: Resolution '{resolution}' not supported by {model} model. "
f"Valid: {', '.join(valid_resolutions)}", file=sys.stderr)
sys.exit(1)
if thinking and model != "nano-banana":
print("Error: --thinking is only supported with Nano Banana model", file=sys.stderr)
sys.exit(1)
if image_grounding and model != "nano-banana":
print("Error: --image-grounding is only supported with Nano Banana model", file=sys.stderr)
sys.exit(1)
# --- Image optimization ---
MAX_DIMENSION = 2048
def optimize_image(img, max_dim=MAX_DIMENSION):
"""Resize if larger than max_dim, preserving aspect ratio."""
from PIL import Image
width, height = img.size
if max(width, height) <= max_dim:
return img
scale = max_dim / max(width, height)
new_size = (round(width * scale), round(height * scale))
return img.resize(new_size, Image.Resampling.LANCZOS)
# --- Prompt logging ---
def save_prompt_log(
log_path: Path,
prompt: str,
output_images: list[Path],
source_images: list[str] | None = None,
model: str | None = None,
):
"""Save the prompt used to generate images as a single .md file."""
timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
content = "# Image Generation Log\n\n"
content += f"**Generated**: {timestamp}\n\n"
if model:
content += f"**Model**: {model}\n\n"
if len(output_images) == 1:
content += f"**Output**: `{output_images[0].name}`\n\n"
else:
content += "**Outputs**:\n"
for img in output_images:
content += f"- `{img.name}`\n"
content += "\n"
if source_images:
content += "**Source Images**:\n"
for src in source_images:
content += f"- `{src}`\n"
content += "\n"
content += f"## Prompt\n\n```\n{prompt}\n```\n"
log_path.write_text(content)
# --- Image extraction ---
def extract_and_save_image(response, output_path: Path,
output_format: str = "png",
quality: int | None = None) -> str | None:
"""Extract image from response and save in requested format.
The API returns images in its preferred format (usually JPEG). We use PIL
to convert to the user's requested format and handle RGBA-to-RGB conversion.
"""
from PIL import Image as PILImage
parts = response.parts if hasattr(response, 'parts') else response.candidates[0].content.parts
text_response = None
image_saved = False
for part in parts:
# Skip thought parts (intermediate reasoning images)
if getattr(part, 'thought', None):
continue
if part.text is not None:
text_response = part.text
elif part.inline_data is not None:
img = PILImage.open(BytesIO(part.inline_data.data))
# Convert RGBA to RGB (compositing alpha onto white)
if img.mode == 'RGBA':
rgb = PILImage.new('RGB', img.size, (255, 255, 255))
rgb.paste(img, mask=img.split()[3])
img = rgb
elif img.mode != 'RGB':
img = img.convert('RGB')
# Save in requested format
if output_format == "jpeg":
save_kwargs = {"format": "JPEG"}
if quality is not None:
save_kwargs["quality"] = quality
img.save(str(output_path), **save_kwargs)
else:
img.save(str(output_path), "PNG")
image_saved = True
if not image_saved:
raise RuntimeError("No image was generated. Check your prompt and try again.")
return text_response
# --- JSON output ---
def format_json_output(results: list[dict], errors: list[dict], total: int) -> str:
"""Format generation results as JSON for agent consumption."""
return json.dumps({
"success": len(errors) == 0,
"generated": len(results),
"total_requested": total,
"images": results,
"errors": errors,
}, indent=2)
# --- Image copying for thread safety ---
def copy_images(images: list) -> list | None:
"""Create deep copies of PIL Images to avoid thread-safety issues."""
if not images:
return None
return [img.copy() for img in images]
# --- Async generation pipeline ---
async def generate_image_async(
client,
model: str,
prompt: str,
output_path: Path,
input_images: list | None = None,
aspect_ratio: str | None = None,
resolution: str | None = None,
grounding: bool = False,
image_grounding: bool = False,
thinking_level: str | None = None,
output_format: str = "png",
quality: int | None = None,
) -> str | None:
"""Generate or edit an image asynchronously."""
from google.genai import types
# Auto-detect resolution if not specified
if resolution is None:
resolution = detect_resolution(input_images or [])
# Build contents: images first (if any), then prompt
if input_images:
contents = input_images + [prompt]
else:
contents = [prompt]
# Build config
config_kwargs = {"response_modalities": ["TEXT", "IMAGE"]}
# Image config
image_config_kwargs = {}
if aspect_ratio:
image_config_kwargs["aspect_ratio"] = aspect_ratio
if resolution:
image_config_kwargs["image_size"] = resolution
if image_config_kwargs:
config_kwargs["image_config"] = types.ImageConfig(**image_config_kwargs)
# Grounding tools
if grounding or image_grounding:
search_types_kwargs = {}
if grounding:
search_types_kwargs["web_search"] = types.WebSearch()
if image_grounding:
search_types_kwargs["image_search"] = types.ImageSearch()
config_kwargs["tools"] = [types.Tool(
google_search=types.GoogleSearch(
search_types=types.SearchTypes(**search_types_kwargs)
)
)]
# Thinking config (Nano Banana only)
if thinking_level and model == "nano-banana":
level_map = {
"minimal": types.ThinkingLevel.MINIMAL,
"low": types.ThinkingLevel.LOW,
"medium": types.ThinkingLevel.MEDIUM,
"high": types.ThinkingLevel.HIGH,
}
config_kwargs["thinking_config"] = types.ThinkingConfig(
thinking_level=level_map[thinking_level]
)
config = types.GenerateContentConfig(**config_kwargs)
# Use async API
response = await client.aio.models.generate_content(
model=MODELS[model],
contents=contents,
config=config,
)
text_response = extract_and_save_image(response, output_path, output_format, quality)
return text_response
async def generate_single(
client,
model: str,
idx: int,
total: int,
out_path: Path,
prompt: str,
input_images: list | None,
aspect_ratio: str | None,
resolution: str | None,
grounding: bool,
image_grounding: bool,
thinking_level: str | None,
output_format: str,
quality: int | None,
) -> tuple[int, Path, str | None, Exception | None]:
"""Generate a single image, return (index, path, text, error)."""
try:
task_images = copy_images(input_images)
text = await generate_image_async(
client=client,
model=model,
prompt=prompt,
output_path=out_path,
input_images=task_images,
aspect_ratio=aspect_ratio,
resolution=resolution,
grounding=grounding,
image_grounding=image_grounding,
thinking_level=thinking_level,
output_format=output_format,
quality=quality,
)
return (idx, out_path, text, None)
except Exception as e:
return (idx, out_path, None, e)
async def run_batch(
client,
model: str,
output_paths: list[Path],
prompt: str,
input_images: list | None,
aspect_ratio: str | None,
resolution: str | None,
grounding: bool,
image_grounding: bool,
thinking_level: str | None,
output_format: str,
quality: int | None,
) -> list[tuple[int, Path, str | None, Exception | None]]:
"""Run batch generation using asyncio.gather for true async parallelism."""
total = len(output_paths)
tasks = [
generate_single(
client=client,
model=model,
idx=i,
total=total,
out_path=path,
prompt=prompt,
input_images=input_images,
aspect_ratio=aspect_ratio,
resolution=resolution,
grounding=grounding,
image_grounding=image_grounding,
thinking_level=thinking_level,
output_format=output_format,
quality=quality,
)
for i, path in enumerate(output_paths, 1)
]
return await asyncio.gather(*tasks)
async def async_main(args, input_images, input_paths, output_paths):
"""Async entry point for image generation."""
from google import genai
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
if not args.quiet:
print("Generating...")
batch_results = await run_batch(
client=client,
model=args.model,
output_paths=output_paths,
prompt=args.prompt,
input_images=input_images,
aspect_ratio=args.aspect,
resolution=args.resolution,
grounding=args.grounding,
image_grounding=args.image_grounding,
thinking_level=args.thinking,
output_format=args.format,
quality=args.quality,
)
# Process results
json_results = []
json_errors = []
results = []
for idx, out_path, text, error in sorted(batch_results, key=lambda x: x[0]):
if error:
if not args.quiet:
print(f"\n[{idx}/{args.batch}] Error: {error}", file=sys.stderr)
json_errors.append({"index": idx, "error": str(error)})
else:
full_path = out_path.resolve()
if not args.quiet:
print(f"\n[{idx}/{args.batch}] Image saved: {full_path}")
# MEDIA line always prints — agents rely on this for image display
print(f"MEDIA: {full_path}")
if text and not args.quiet:
print(f"Model response: {text}")
results.append(full_path)
json_results.append({
"index": idx,
"path": str(full_path),
"model_response": text,
})
if args.json:
print(format_json_output(json_results, json_errors, args.batch))
return results
# --- CLI ---
def main():
parser = argparse.ArgumentParser(
description="Generate and edit images using Gemini Image API",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__
)
parser.add_argument(
"--prompt", "-p",
required=True,
help="Image description or edit instruction"
)
parser.add_argument(
"--output", "-o",
required=True,
help="Output file path (e.g., output.png)"
)
parser.add_argument(
"--input", "-i",
action="append",
dest="inputs",
help="Input image path for editing/composition (can be repeated up to 14 times)"
)
parser.add_argument(
"--model", "-m",
choices=["nano-banana", "pro"],
default=DEFAULT_MODEL,
help="Model: nano-banana (default) or pro"
)
parser.add_argument(
"--aspect", "-a",
help="Aspect ratio (e.g., 1:1, 16:9, 9:16)"
)
parser.add_argument(
"--resolution", "-r",
help="Output resolution: 0.5K, 1K, 2K, 4K (default: auto-detect or 1K)"
)
parser.add_argument(
"--grounding", "-g",
action="store_true",
help="Enable Google Search web grounding"
)
parser.add_argument(
"--image-grounding",
action="store_true",
help="Enable image search grounding (Nano Banana only, use with --grounding)"
)
parser.add_argument(
"--thinking", "-t",
choices=["minimal", "low", "medium", "high"],
default=None,
help="Thinking level (Nano Banana only, default: minimal)"
)
parser.add_argument(
"--quality", "-q",
type=int,
default=None,
help="Output compression quality 1-100 (JPEG only)"
)
parser.add_argument(
"--format", "-f",
choices=["png", "jpeg"],
default="png",
help="Output image format (default: png)"
)
parser.add_argument(
"--batch", "-b",
type=int,
choices=[1, 2, 3, 4],
default=1,
help="Generate multiple variations (1-4, default: 1)"
)
parser.add_argument(
"--json",
action="store_true",
help="Output results as JSON (for agent consumption)"
)
parser.add_argument(
"--quiet",
action="store_true",
help="Suppress progress output (MEDIA lines still printed)"
)
args = parser.parse_args()
# Check API key early, before any output
if not os.environ.get("GEMINI_API_KEY"):
print("Error: GEMINI_API_KEY environment variable not set", file=sys.stderr)
sys.exit(2)
# Validate input count
if args.inputs and len(args.inputs) > 14:
print("Error: Maximum 14 input images allowed", file=sys.stderr)
sys.exit(1)
# Validate quality range
if args.quality is not None:
if args.quality < 1 or args.quality > 100:
print("Error: --quality must be between 1 and 100", file=sys.stderr)
sys.exit(1)
if args.format != "jpeg":
print("Warning: --quality only affects JPEG output, ignored for PNG", file=sys.stderr)
# Validate model-specific parameters
validate_model_params(args.model, args.aspect, args.resolution,
args.thinking, args.image_grounding)
# Set up output path with correct extension
output_path = Path(args.output)
if args.format == "jpeg" and output_path.suffix.lower() not in (".jpg", ".jpeg"):
output_path = output_path.with_suffix(".jpg")
elif args.format == "png" and output_path.suffix.lower() != ".png":
output_path = output_path.with_suffix(".png")
if not output_path.parent.exists():
output_path.parent.mkdir(parents=True, exist_ok=True)
# Load input images if provided
input_images = None
input_paths = []
if args.inputs:
from PIL import Image
input_images = []
for img_path in args.inputs:
try:
img = Image.open(img_path)
img.load()
original_size = img.size
img = optimize_image(img)
input_images.append(img)
input_paths.append(img_path)
if not args.quiet:
if img.size != original_size:
print(f"Loaded: {img_path} ({original_size[0]}x{original_size[1]} → {img.size[0]}x{img.size[1]})")
else:
print(f"Loaded: {img_path} ({img.size[0]}x{img.size[1]})")
except Exception as e:
print(f"Error loading {img_path}: {e}", file=sys.stderr)
sys.exit(1)
# Auto-detect aspect ratio from base image (first input) if not specified
if args.aspect:
if not args.quiet:
print(f"Aspect ratio: {args.aspect}")
elif input_images:
base_img = input_images[0]
args.aspect = get_closest_aspect_ratio(base_img.width, base_img.height, model=args.model)
if not args.quiet:
print(f"Auto aspect ratio: {args.aspect} (from base image)")
else:
args.aspect = "1:1"
if not args.quiet:
print(f"Auto aspect ratio: {args.aspect} (default)")
# Determine mode for display
if not input_images:
mode = "t2i (text-to-image)"
elif len(input_images) == 1:
mode = "i2i (image editing)"
else:
mode = f"multi-reference ({len(input_images)} images)"
resolution = args.resolution or detect_resolution(input_images or [])
model_display = MODEL_DISPLAY[args.model]
if not args.quiet:
print(f"Model: {model_display}")
print(f"Mode: {mode}")
print(f"Resolution: {resolution}")
if args.grounding:
print("Web search grounding: enabled")
if args.image_grounding:
print("Image search grounding: enabled")
if args.thinking:
print(f"Thinking: {args.thinking}")
if args.batch > 1:
print(f"Batch: {args.batch} images (async parallel)")
# Generate output paths for batch
if args.batch == 1:
output_paths = [output_path]
else:
stem = output_path.stem
suffix = output_path.suffix
parent = output_path.parent
output_paths = [parent / f"{stem}-{i}{suffix}" for i in range(1, args.batch + 1)]
# Run async main
results = asyncio.run(async_main(args, input_images, input_paths, output_paths))
if not results:
print("Error: No images were generated", file=sys.stderr)
sys.exit(1)
# Save single prompt log for all generated images
log_path = output_path.with_suffix(".md")
save_prompt_log(log_path, args.prompt, results,
input_paths if input_paths else None,
model=MODELS[args.model])
if not args.quiet:
print(f"\nPrompt log: {log_path.resolve()}")
print(f"Generated {len(results)}/{args.batch} images")
if __name__ == "__main__":
main()
Related skills
How it compares
Prompt-pattern skill for generative art—not a hosted image API or design-system component library.
FAQ
Who is generate-image for?
Developers and designers using AI coding agents who need stronger image-generation prompts for marketing, product, and UI collateral during product build.
When should I use generate-image?
In the Build phase while crafting prompts for heroes, product shots, stickers, or branded text renders before launch creative goes live.
Is generate-image safe to install?
Check the Security Audits panel on this Prism page; the skill is instructional text only but your image tool may need network/API access separately.