
Nano Banana
- 15 installs
- 15 repo stars
- Updated August 1, 2026
- connorads/dotfiles
Generates and edits images with Google's Gemini Nano Banana models, supporting text-to-image, up to 14 reference images, and 0.5K-4K resolution.
About
Generates and edits images using Google's Gemini image models (Nano Banana 2 default, Pro legacy) via a script, supporting text-to-image and image editing with reference images. A developer uses it to create or modify images with configurable resolution and aspect ratio.
- Text-to-image and image editing via Google Gemini image models (Nano Banana 2 default)
- Supports up to 14 reference images, 0.5K-4K resolution, aspect ratio, and adjustable thinking
Nano Banana by the numbers
- 15 all-time installs (skills.sh)
- Ranked #1,028 of 1,337 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/connorads/dotfiles --skill nano-bananaAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 15 |
|---|---|
| repo stars | ★ 15 |
| Last updated | August 1, 2026 |
| Repository | connorads/dotfiles ↗ |
What it does
Generates and edits images with Google's Gemini Nano Banana models, supporting text-to-image, up to 14 reference images, and 0.5K-4K resolution.
Files
Nano Banana Image Generation & Editing
Generate new images or edit existing ones using Google's Gemini image API.
Default model: Nano Banana 2 (gemini-3.1-flash-image-preview) — cheaper, faster, more features. Legacy model: Nano Banana Pro (gemini-3-pro-image-preview) — available via --model nano-banana-pro, deprecated 9 March 2026.
Usage
Run the script using absolute path (do NOT cd to skill directory first):
Generate new image:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "your image description" --filename "output.png" [--resolution 0.5K|1K|2K|4K] [--aspect-ratio 16:9] [--thinking high] [--api-key KEY]Edit existing image:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "editing instructions" --filename "output.png" --input-image "path/to/input.png" [--resolution 0.5K|1K|2K|4K] [--api-key KEY]Multi-reference generation (up to 14 images):
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "combine these styles" --filename "output.png" --input-image "ref1.png" --input-image "ref2.png"Important: Always run from the user's current working directory so images are saved where the user is working, not in the skill directory.
Model Selection
| Flag value | Model ID | Notes |
|---|---|---|
nano-banana-2 (default) | gemini-3.1-flash-image-preview | Cheaper, faster, 0.5K support, aspect ratio, thinking, 14 ref images |
nano-banana-pro | gemini-3-pro-image-preview | Deprecated 9 March 2026 |
Resolution Options
Uppercase K required:
- 0.5K — ~512px (cheapest, Nano Banana 2 only)
- 1K (default) — ~1024px
- 2K — ~2048px
- 4K — ~4096px
Map user requests:
- "thumbnail", "small", "cheap", "low resolution" →
0.5K - No mention / "1080", "1080p", "1K" →
1K - "2K", "2048", "medium resolution" →
2K - "high resolution", "high-res", "hi-res", "4K", "ultra" →
4K
When editing, resolution auto-detects from input image size if not specified.
Aspect Ratio
14 supported ratios (Nano Banana 2 only). Omit to let the API decide.
Map user requests:
- "landscape", "wide" →
16:9 - "portrait", "phone", "story", "reel" →
9:16 - "square" →
1:1 - "cinematic", "ultrawide" →
21:9 - "photo portrait" →
3:4or2:3
Thinking
Controls reasoning effort for complex prompts (Nano Banana 2 only). Omit for fastest/cheapest.
--thinking minimal— slightly better quality, low latency impact--thinking high— best for complex composition, text rendering, detailed scenes
Note: thinking tokens are billed regardless of level.
Reference Images
Up to 14 input images via repeated --input-image flags. Use cases:
- Style transfer — provide style reference + subject
- Character consistency — provide character reference images
- Object reference — provide object images to include
- Image editing — provide single image + editing instructions in prompt
Resolution auto-detects from the first input image when --resolution not specified.
API Key
Checked in order: 1. --api-key argument 2. GEMINI_API_KEY environment variable
Filename Generation
Pattern: yyyy-mm-dd-hh-mm-ss-name.png
- Timestamp: current date/time, 24-hour format
- Name: descriptive lowercase with hyphens (1-5 words)
- Unclear context: use random identifier (e.g.,
x9k2)
Examples:
- "A serene Japanese garden" →
2026-03-01-14-23-05-japanese-garden.png - Unclear →
2026-03-01-17-12-48-x9k2.png
Image Editing
1. Check if user provides an image path or references an image in the current directory 2. Use --input-image with the path 3. Pass editing instructions in --prompt 4. Common tasks: add/remove elements, change style, adjust colours, blur background
Prompt Handling
For generation: Pass user's image description as-is to --prompt. Only rework if clearly insufficient.
For editing: Pass editing instructions in --prompt (e.g., "add a rainbow in the sky", "make it look like a watercolour painting")
Preserve user's creative intent in both cases.
Output
- Saves PNG to current directory (or specified path if filename includes directory)
- Script outputs the full path to the generated image
- Do not read the image back — just inform the user of the saved path
Pricing (Paid Tier Only)
No free tier for image generation.
Nano Banana 2 (gemini-3.1-flash-image-preview) — default
| Resolution | Standard | Batch |
|---|---|---|
| 0.5K (512px) | $0.045 | $0.022 |
| 1K | $0.067 | $0.034 |
| 2K | $0.101 | $0.050 |
| 4K | $0.151 | $0.076 |
Google Search grounding: 5,000 prompts/month free, then $14/1,000 queries.
Nano Banana Pro (gemini-3-pro-image-preview) — deprecated 9 Mar 2026
| Component | Standard | Batch |
|---|---|---|
| Input (per image) | $0.0011 | $0.0006 |
| Output 1K/2K | $0.134 | $0.067 |
| Output 4K | $0.24 | $0.12 |
Examples
Generate with aspect ratio:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "Cinematic landscape at golden hour" --filename "2026-03-01-10-00-00-landscape.png" --resolution 2K --aspect-ratio 21:9Generate with thinking for complex prompt:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "A detailed infographic about climate change with text labels" --filename "2026-03-01-10-05-00-infographic.png" --thinking high --resolution 4KMulti-reference style transfer:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "A portrait in this artistic style" --filename "2026-03-01-10-10-00-styled-portrait.png" --input-image "style-ref.png" --input-image "subject.png"Budget thumbnail:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "Simple icon of a house" --filename "2026-03-01-10-15-00-house-icon.png" --resolution 0.5KEdit existing image:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "make the sky more dramatic with storm clouds" --filename "2026-03-01-14-25-30-dramatic-sky.png" --input-image "original-photo.jpg" --resolution 2KUse legacy model:
uv run ~/.claude/skills/nano-banana/scripts/generate_image.py --prompt "A serene Japanese garden" --filename "2026-03-01-14-23-05-japanese-garden.png" --model nano-banana-pro --resolution 4K#!/usr/bin/env python3
# /// script
# requires-python = ">=3.10"
# dependencies = [
# "google-genai>=1.0.0",
# "pillow>=10.0.0",
# ]
# ///
"""
Generate and edit images using Google's Gemini image generation API.
Supports Nano Banana 2 (gemini-3.1-flash-image-preview, default) and
legacy Nano Banana Pro (gemini-3-pro-image-preview, deprecated 9 Mar 2026).
Usage:
uv run generate_image.py --prompt "description" --filename "out.png" [options]
"""
import argparse
import os
import sys
from pathlib import Path
MODEL_MAP: dict[str, str] = {
"nano-banana-2": "gemini-3.1-flash-image-preview",
"nano-banana-pro": "gemini-3-pro-image-preview", # deprecated 9 Mar 2026
}
ASPECT_RATIOS = [
"1:1", "1:4", "1:8", "2:3", "3:2", "3:4", "4:1",
"4:3", "4:5", "5:4", "8:1", "9:16", "16:9", "21:9",
]
MAX_REFERENCE_IMAGES = 14
def get_api_key(provided_key: str | None) -> str | None:
"""Get API key from argument first, then environment."""
if provided_key:
return provided_key
return os.environ.get("GEMINI_API_KEY")
def auto_detect_resolution(width: int, height: int) -> str:
"""Map input image dimensions to an appropriate output resolution."""
max_dim = max(width, height)
if max_dim >= 3000:
return "4K"
if max_dim >= 1500:
return "2K"
if max_dim >= 800:
return "1K"
return "0.5K"
def main():
parser = argparse.ArgumentParser(
description="Generate/edit images using Google Gemini (Nano Banana)"
)
parser.add_argument(
"--prompt", "-p",
required=True,
help="Image description or editing instructions",
)
parser.add_argument(
"--filename", "-f",
required=True,
help="Output filename (e.g., sunset-mountains.png)",
)
parser.add_argument(
"--input-image", "-i",
action="append",
default=[],
help="Input image path for editing/reference (repeat up to 14 times)",
)
parser.add_argument(
"--model", "-m",
choices=list(MODEL_MAP),
default="nano-banana-2",
help="Model: nano-banana-2 (default) or nano-banana-pro (deprecated 9 Mar 2026)",
)
parser.add_argument(
"--resolution", "-r",
choices=["0.5K", "1K", "2K", "4K"],
default="1K",
help="Output resolution: 0.5K, 1K (default), 2K, or 4K",
)
parser.add_argument(
"--aspect-ratio", "-a",
choices=ASPECT_RATIOS,
default=None,
help="Output aspect ratio (e.g., 16:9, 9:16, 1:1). Default: API decides.",
)
parser.add_argument(
"--thinking",
choices=["minimal", "high"],
default=None,
help="Thinking level for complex prompts (Nano Banana 2 only)",
)
parser.add_argument(
"--api-key", "-k",
help="Gemini API key (overrides GEMINI_API_KEY env var)",
)
args = parser.parse_args()
# Validate reference image count
if len(args.input_image) > MAX_REFERENCE_IMAGES:
print(f"Error: Maximum {MAX_REFERENCE_IMAGES} reference images supported.", file=sys.stderr)
sys.exit(1)
# Get API key
api_key = get_api_key(args.api_key)
if not api_key:
print("Error: No API key provided.", file=sys.stderr)
print("Please either:", file=sys.stderr)
print(" 1. Provide --api-key argument", file=sys.stderr)
print(" 2. Set GEMINI_API_KEY environment variable", file=sys.stderr)
sys.exit(1)
# Import here after checking API key to avoid slow import on error
from google import genai
from google.genai import types
from PIL import Image as PILImage
# Resolve model
model_id = MODEL_MAP[args.model]
# Initialise client
client = genai.Client(api_key=api_key)
# Set up output path
output_path = Path(args.filename)
output_path.parent.mkdir(parents=True, exist_ok=True)
# Load input images
input_images: list[PILImage.Image] = []
output_resolution = args.resolution
for image_path in args.input_image:
try:
img = PILImage.open(image_path)
input_images.append(img)
print(f"Loaded input image: {image_path}")
except Exception as e:
print(f"Error loading input image '{image_path}': {e}", file=sys.stderr)
sys.exit(1)
# Auto-detect resolution from first input image if user didn't set it
if input_images and args.resolution == "1K":
width, height = input_images[0].size
output_resolution = auto_detect_resolution(width, height)
print(f"Auto-detected resolution: {output_resolution} (from input {width}x{height})")
# Build contents (images first if editing, prompt only if generating)
if input_images:
contents: str | list[PILImage.Image | str] = [*input_images, args.prompt]
print(f"Editing image with resolution {output_resolution}...")
else:
contents = args.prompt
print(f"Generating image with resolution {output_resolution}...")
# Build config dynamically
image_config_kwargs: dict[str, str] = {"image_size": output_resolution}
if args.aspect_ratio:
image_config_kwargs["aspect_ratio"] = args.aspect_ratio
config_kwargs: dict[str, object] = {
"response_modalities": ["TEXT", "IMAGE"],
"image_config": types.ImageConfig(**image_config_kwargs),
}
if args.thinking and model_id == MODEL_MAP["nano-banana-2"]:
config_kwargs["thinking_config"] = types.ThinkingConfig(
thinking_level=args.thinking,
include_thoughts=True,
)
try:
response = client.models.generate_content(
model=model_id,
contents=contents,
config=types.GenerateContentConfig(**config_kwargs),
)
# Process response and convert to PNG
image_saved = False
for part in response.parts:
if hasattr(part, "thought") and part.thought:
print(f"Thinking: {part.text}")
elif part.text is not None:
print(f"Model response: {part.text}")
elif part.inline_data is not None:
# Convert inline data to PIL Image and save as PNG
from io import BytesIO
# inline_data.data is already bytes, not base64
image_data = part.inline_data.data
if isinstance(image_data, str):
import base64
image_data = base64.b64decode(image_data)
image = PILImage.open(BytesIO(image_data))
# Ensure RGB mode for PNG
if image.mode == "RGBA":
rgb_image = PILImage.new("RGB", image.size, (255, 255, 255))
rgb_image.paste(image, mask=image.split()[3])
rgb_image.save(str(output_path), "PNG")
elif image.mode == "RGB":
image.save(str(output_path), "PNG")
else:
image.convert("RGB").save(str(output_path), "PNG")
image_saved = True
if image_saved:
full_path = output_path.resolve()
print(f"\nImage saved: {full_path}")
else:
print("Error: No image was generated in the response.", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error generating image: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()