
Imagen
- 507 installs
- 364 repo stars
- Updated July 9, 2026
- sanjay3290/ai-skills
imagen is a Claude Code skill that generates images through Google Gemini's gemini-3-pro-image-preview model for developers who need UI mockups, icons, diagrams, and visual assets inside coding sessions.
About
imagen is version 1.0 of a sanjay3290/ai-skills Apache-2.0 skill that creates images using Google Gemini's `gemini-3-pro-image-preview` image generation model during Claude Code sessions. Developers invoke imagen for UI mockups, icons, illustrations, diagrams, concept art, or placeholder visuals while building frontends or writing documentation. The skill removes context switching to external design tools for quick asset iteration. Reach for imagen when a session needs generated visuals tied to feature work rather than stock photo searches or manual Figma exports.
- Generates images using gemini-3-pro-image-preview model
- Saves output directly to your project directory with returned file path
- Cross-platform support for Windows, macOS, and Linux
- Activates automatically on prompts like "generate an image of..." or "create a picture..."
- Supports UI mockups, icons, illustrations, diagrams, concept art, and placeholders
Imagen by the numbers
- 507 all-time installs (skills.sh)
- +12 installs in the week ending Jul 26, 2026 (Skillselion tracking)
- Ranked #402 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/sanjay3290/ai-skills --skill imagenAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 507 |
|---|---|
| repo stars | ★ 364 |
| Last updated | July 9, 2026 |
| Repository | sanjay3290/ai-skills ↗ |
How do you generate images inside Claude Code?
Generate images on demand directly inside Claude Code sessions for UI mockups, diagrams, icons, and visual assets.
Who is it for?
Developers in Claude Code who need fast Gemini-generated visuals for UI mockups, docs diagrams, or placeholder assets without leaving the editor.
Skip if: Production brand asset pipelines requiring strict style guides, print resolution workflows, or non-Gemini image model requirements.
When should I use this skill?
A developer asks to create, generate, or produce images for mockups, icons, illustrations, diagrams, or visual placeholders.
What you get
Generated image files for mockups, icons, illustrations, diagrams, or concept art saved from the Claude session.
- generated image files
- ui mockup or diagram assets
By the numbers
- Skill version 1.0 under Apache-2.0 license
- Uses gemini-3-pro-image-preview image generation model
Files
Imagen - AI Image Generation Skill
Overview
This skill generates images using Google Gemini's image generation model (gemini-3-pro-image-preview). It enables seamless image creation during any Claude Code session - whether you're building frontend UIs, creating documentation, or need visual representations of concepts.
Cross-Platform: Works on Windows, macOS, and Linux.
When to Use This Skill
Automatically activate this skill when:
- User requests image generation (e.g., "generate an image of...", "create a picture...")
- Frontend development requires placeholder or actual images
- Documentation needs illustrations or diagrams
- Visualizing concepts, architectures, or ideas
- Creating icons, logos, or UI assets
- Any task where an AI-generated image would be helpful
How It Works
1. Takes a text prompt describing the desired image 2. Calls Google Gemini API with image generation configuration 3. Saves the generated image to a specified location (defaults to current directory) 4. Returns the file path for use in your project
Usage
Python (Cross-Platform - Recommended)
# Basic usage
python scripts/generate_image.py "A futuristic city skyline at sunset"
# With custom output path
python scripts/generate_image.py "A minimalist app icon for a music player" "./assets/icons/music-icon.png"
# With custom size
python scripts/generate_image.py --size 2K "High resolution landscape" "./wallpaper.png"Requirements
GEMINI_API_KEYenvironment variable must be set- Python 3.6+ (uses standard library only, no pip install needed)
Output
Generated images are saved as PNG files. The script returns:
- Success: Path to the generated image
- Failure: Error message with details
Examples
Frontend Development
User: "I need a hero image for my landing page - something abstract and tech-focused"
-> Generates and saves image, provides path for use in HTML/CSSDocumentation
User: "Create a diagram showing microservices architecture"
-> Generates visual representation, ready for README or docsUI Assets
User: "Generate a placeholder avatar image for the user profile component"
-> Creates image in appropriate size for component use# Google Gemini API Key
# Get your free key at: https://aistudio.google.com/
GEMINI_API_KEY=your-api-key-here
# Optional: Default image size (512, 1K, or 2K)
# IMAGE_SIZE=1K
# Environment files with credentials
.env
# Generated images (user should manage their own)
*.png
*.jpg
*.jpeg
# Python
__pycache__/
*.py[cod]
*$py.class
.Python
*.so
# Virtual environments
venv/
ENV/
env/
.venv/
# IDE
.idea/
.vscode/
*.swp
*.swo
*~
# OS
.DS_Store
Thumbs.db
Imagen Skill - Usage Examples
Frontend Development Scenarios
Hero Image for Landing Page
User: "I need a hero image for my fintech startup landing page"
Claude will run:
python scripts/generate_image.py "Modern fintech hero image with abstract flowing gradients in blue and purple, suggesting financial growth and technology, clean minimalist style" "./src/assets/hero.png"App Icons
User: "Generate an app icon for my weather app"
Claude will run:
python scripts/generate_image.py --size 512 "Minimalist weather app icon, stylized sun partially behind a cloud, warm orange and blue gradient, rounded square suitable for iOS/Android" "./assets/icons/weather-icon.png"Placeholder Images
User: "I need placeholder images for my e-commerce product grid"
Claude will run:
python scripts/generate_image.py "E-commerce product placeholder image, neutral gray background, subtle product box outline, professional and clean" "./public/images/placeholder-product.png"Background Patterns
User: "Create a subtle background pattern for my website"
Claude will run:
python scripts/generate_image.py --size 2K "Seamless tileable subtle geometric pattern, very light gray on white, minimalist, suitable for website background" "./assets/bg-pattern.png"Documentation & Diagrams
Architecture Diagram
User: "Visualize our microservices architecture"
Claude will run:
python scripts/generate_image.py "Clean technical diagram showing microservices architecture, boxes connected by arrows, services labeled API Gateway, Auth Service, User Service, Order Service, white background, professional style" "./docs/images/architecture.png"Process Flowchart
User: "Create a flowchart for user onboarding"
Claude will run:
python scripts/generate_image.py "Flowchart diagram showing user onboarding process, start with signup, branches for email verification, profile setup, tutorial completion, clean business style with blue accents" "./docs/onboarding-flow.png"Marketing & Creative
Social Media Graphics
User: "I need a banner for our product launch on Twitter"
Claude will run:
python scripts/generate_image.py "Modern product launch banner 1200x675, bold typography space for text overlay, dynamic abstract tech background, vibrant gradients purple to blue" "./marketing/twitter-banner.png"Blog Post Headers
User: "Generate a header image for my blog post about AI"
Claude will run:
python scripts/generate_image.py "Blog header image about artificial intelligence, abstract neural network visualization, glowing nodes and connections, dark background with blue accent lights, futuristic and professional" "./blog/images/ai-header.png"Professional Photography (Advanced Techniques)
Corporate Headshots
User: "Generate a professional corporate headshot for our team page"
Claude will run:
python scripts/generate_image.py "Professional corporate headshot, 85mm f/2.8 lens, three-point studio lighting with soft key light from left, subtle fill light, rim light for hair separation, navy blue business suit, neutral gray backdrop, visible catchlights in eyes, chest-up framing, natural skin texture with subtle pores, sharp focus on eyes" "./team/headshot-template.png"Casual Lifestyle Portraits
User: "Create a casual, authentic-looking portrait for my personal brand"
Claude will run:
python scripts/generate_image.py "Casual lifestyle portrait, Kodak Portra 400 color tones, natural window light with soft shadows, 3/4 body shot, authentic film grain, early-2000s digital camera aesthetic, relaxed pose, natural skin texture, warm highlights, slight background blur" "./branding/casual-portrait.png"High-Fashion Editorial
User: "Generate a high-fashion editorial style photo"
Claude will run:
python scripts/generate_image.py "High-fashion editorial portrait, dramatic studio lighting with strong key light, deep shadows, 85mm f/1.4 lens with shallow depth of field, model wearing designer clothing, fabric texture visible, silk sheen and drape, backstage atmosphere, bold composition, magazine-quality, keep facial features exactly consistent" "./fashion/editorial-shot.png"Product Photography with Model
User: "Create a lifestyle product shot showing someone using headphones"
Claude will run:
python scripts/generate_image.py "Lifestyle product photography, person wearing premium over-ear headphones, 50mm prime lens, natural soft lighting from large window, clean modern interior background with bokeh, 3/4 profile view, natural skin texture, subtle catchlights, product details sharp and visible, authentic lifestyle feel, Fujifilm Pro 400H color tones" "./products/headphones-lifestyle.png"Social Media Selfie Style
User: "Generate an authentic-looking selfie for social media content"
Claude will run:
python scripts/generate_image.py "Authentic social media selfie, iPhone camera quality with realistic sensor noise, bathroom mirror selfie angle, natural indoor lighting, slight lens distortion, casual outfit, natural skin with subtle imperfections, relaxed authentic expression, 1990s disposable camera aesthetic, do not alter the face, early morning soft light" "./content/selfie-style.png"E-commerce Fashion
User: "Create a product photo for an e-commerce clothing listing"
Claude will run:
python scripts/generate_image.py "E-commerce fashion photography, model wearing casual cotton t-shirt, clean white studio background, soft even lighting, full body shot, fabric texture clearly visible, natural pose, 70mm lens, product colors accurate, subtle shadows for depth, professional catalog style, natural skin texture" "./ecommerce/tshirt-listing.png"Command Line Usage
# Generate with default settings
python scripts/generate_image.py "A peaceful zen garden"
# Specify output location
python scripts/generate_image.py "Mountain landscape" "./wallpapers/mountains.png"
# High resolution output
python scripts/generate_image.py --size 2K "Detailed cityscape" "./high-res/city.png"
# Quick thumbnail/icon
python scripts/generate_image.py --size 512 "Simple checkmark icon" "./icons/check.png"Windows PowerShell
# Generate with default settings
python scripts/generate_image.py "A peaceful zen garden"
# With environment variable for size
$env:IMAGE_SIZE = "2K"
python scripts/generate_image.py "Detailed cityscape" "./high-res/city.png"Batch Generation (via shell)
macOS/Linux:
# Generate multiple variations
for i in 1 2 3; do
python scripts/generate_image.py "Abstract art variation $i" "./art/abstract-$i.png"
doneWindows PowerShell:
# Generate multiple variations
1..3 | ForEach-Object {
python scripts/generate_image.py "Abstract art variation $_" "./art/abstract-$_.png"
}Integration with Frontend Workflow
React Project
User: "Add a loading spinner image to my React components"
Claude will:
1. Generate the image: python scripts/generate_image.py "Minimal loading spinner, circular design, gradient from light to dark blue, clean vector style" "./src/assets/spinner.png"
2. Update component to use the new imageVue Project
User: "Need an empty state illustration for when there's no data"
Claude will:
1. Generate: python scripts/generate_image.py "Friendly empty state illustration, person looking at empty box, soft pastel colors, friendly and approachable, modern SaaS style" "./src/assets/empty-state.png"
2. Add to the component templateFlutter Project
User: "Generate an onboarding illustration"
Claude will:
1. Generate: python scripts/generate_image.py --size 2K "Mobile app onboarding illustration, person using smartphone, abstract flowing shapes in background, modern gradient style" "./assets/images/onboarding.png"
2. Add to pubspec.yaml assets
3. Use in the onboarding widgetTips for Better Results
Basic Tips
1. Be Descriptive: The more detail, the better the output 2. Mention Style: "flat design", "3D render", "watercolor", "minimalist" 3. Specify Colors: "blue and orange gradient", "monochrome", "pastel colors" 4. Include Context: "suitable for dark mode", "web-safe", "mobile app" 5. Define Purpose: "for hero section", "app icon", "documentation"
Advanced Photography Tips
6. Specify Camera Settings: "85mm f/1.4 lens", "50mm prime", "shallow depth of field" 7. Define Lighting Setup: "three-point lighting", "soft key light from left", "rim light for separation" 8. Reference Film Stocks: "Kodak Portra 400 tones", "Fujifilm Pro 400H", "subtle film grain" 9. Add Era Aesthetics: "1990s disposable camera quality", "early-2000s digital look" 10. Include Texture Details: "natural skin texture with visible pores", "fabric weave visible" 11. Specify Framing: "chest-up framing", "3/4 body shot", "rule of thirds placement" 12. Preserve Faces: When editing, add "keep facial features exactly consistent" or "do not alter the face"
Imagen - AI Image Generation Skill
A Claude Code skill that generates images using Google Gemini's image generation model. Simply ask Claude to create an image during any coding session, and it will generate and save it for you.
Cross-Platform: Works on Windows, macOS, and Linux.
Quick Start
1. Install the Skill
# Via plugin marketplace (recommended)
/plugin install imagen@ai-skills
# Or add the entire marketplace
/plugin marketplace add sanjay3290/ai-skills2. Set Your API Key
Get a free API key from Google AI Studio:
Windows (PowerShell):
$env:GEMINI_API_KEY = "your-api-key-here"Windows (CMD):
set GEMINI_API_KEY=your-api-key-heremacOS/Linux:
export GEMINI_API_KEY="your-api-key-here"
# Add to ~/.zshrc or ~/.bashrc to persist
echo 'export GEMINI_API_KEY="your-api-key-here"' >> ~/.zshrc3. Use It!
Just ask Claude to generate an image during any conversation:
"Generate an image of a sunset over mountains"
"I need a hero image for my landing page"
"Create an app icon for a weather app"How It Works
When you mention needing an image, Claude will automatically: 1. Recognize the request and activate this skill 2. Call the Google Gemini API with your prompt 3. Save the generated image to your project 4. Tell you where to find it
Features
- Cross-Platform: Python script works on Windows, macOS, and Linux
- Automatic Activation: Claude detects when you need an image
- Multiple Sizes: 512px, 1K (default), or 2K resolution
- Custom Output Paths: Save images wherever you need them
- Frontend Ready: Perfect for UI development, placeholders, icons
- Documentation Images: Generate diagrams, illustrations, flowcharts
Usage Examples
During Frontend Development
You: "I'm building a dashboard. Generate a placeholder chart image."
Claude: *generates and saves image* "Created the image at ./assets/chart-placeholder.png"For Documentation
You: "Create an architecture diagram for our microservices"
Claude: *generates image* "Saved to ./docs/architecture.png"Custom Size & Location
You: "Generate a high-res hero image and save it to ./public/hero.png"
Claude: *generates 2K image at specified path*Configuration
Environment Variables
| Variable | Required | Default | Description |
|---|---|---|---|
GEMINI_API_KEY | Yes | - | Your Google Gemini API key |
IMAGE_SIZE | No | 1K | Default size (512, 1K, or 2K) |
GEMINI_MODEL | No | gemini-3-pro-image-preview | Gemini model ID |
Image Sizes
| Size | Resolution | Best For |
|---|---|---|
512 | 512x512 | Icons, thumbnails, quick previews |
1K | 1024x1024 | General use, web images |
2K | 2048x2048 | High-res, print, retina displays |
Manual Script Usage
# Basic usage
python scripts/generate_image.py "A serene lake at dawn"
# Custom output path
python scripts/generate_image.py "App icon" "./icon.png"
# With size option
python scripts/generate_image.py --size 2K "Detailed landscape" "./wallpaper.png"
# With custom model
python scripts/generate_image.py --model gemini-3-pro-image-preview "A logo" "./logo.png"File Structure
imagen/
├── SKILL.md # Skill definition (Claude reads this)
├── README.md # This file
├── reference.md # Detailed API reference
├── examples.md # Usage examples
├── .env.example # API key template
└── scripts/
└── generate_image.py # Cross-platform Python scriptRequirements
- Python 3.6+: No additional packages needed (uses standard library only)
- Gemini API Key: Free from Google AI Studio
Troubleshooting
"GEMINI_API_KEY not set"
Windows:
echo $env:GEMINI_API_KEY # Should show your keymacOS/Linux:
echo $GEMINI_API_KEY # Should show your key
source ~/.zshrc # Reload if you just added itAPI Errors
- 400: Check your prompt for special characters
- 429: Rate limited, wait and retry
- 403: Invalid API key
Getting a Gemini API Key
1. Go to Google AI Studio 2. Sign in with your Google account 3. Click "Get API Key" in the left sidebar 4. Create a new key or copy an existing one 5. Add it to your environment as shown above
License
Apache-2.0 License - See LICENSE for details.
Imagen Skill Reference
Setup
1. Get a Gemini API Key
1. Go to Google AI Studio 2. Click "Get API Key" 3. Create a new API key or use an existing one
2. Set Environment Variable
Windows (PowerShell):
$env:GEMINI_API_KEY = "your-api-key-here"
# To persist across sessions, add to your PowerShell profile:
Add-Content $PROFILE "`n`$env:GEMINI_API_KEY = 'your-api-key-here'"Windows (CMD):
set GEMINI_API_KEY=your-api-key-here
# To persist, use System Properties > Environment VariablesmacOS/Linux:
export GEMINI_API_KEY="your-api-key-here"
# Add to ~/.zshrc or ~/.bashrc to persist
echo 'export GEMINI_API_KEY="your-api-key-here"' >> ~/.zshrcAPI Reference
Model
- Model ID:
gemini-3-pro-image-preview(configurable via--modelflag orGEMINI_MODELenv var) - Endpoint:
https://generativelanguage.googleapis.com/v1beta/models/{model}:streamGenerateContent
Image Sizes
| Size | Description |
|---|---|
512 | 512x512 pixels - Fast, good for icons/thumbnails |
1K | 1024x1024 pixels - Default, balanced quality/speed |
2K | 2048x2048 pixels - High resolution, slower |
Script Parameters
Python Script (Cross-Platform)
python scripts/generate_image.py <prompt> [output_path] [--size SIZE]| Parameter | Required | Default | Description |
|---|---|---|---|
prompt | Yes | - | Text description of desired image |
output_path | No | ./generated-image.png | Where to save the image |
--size | No | 1K | Image size (512, 1K, or 2K) |
Environment Variables
| Variable | Required | Default | Description |
|---|---|---|---|
GEMINI_API_KEY | Yes | - | Your Google Gemini API key |
IMAGE_SIZE | No | 1K | Image size (512, 1K, or 2K) |
Usage Examples
Basic Generation
python scripts/generate_image.py "A serene mountain landscape at dawn"Custom Output Path
python scripts/generate_image.py "Minimalist logo design" "./assets/logo.png"High Resolution
python scripts/generate_image.py --size 2K "Detailed portrait" "./high-res.png"Small/Fast Generation
python scripts/generate_image.py --size 512 "Simple icon" "./icon.png"Prompt Tips
For Best Results
1. Be specific: "A red sports car" vs "A cherry red 1967 Mustang convertible" 2. Include style: "in watercolor style", "photorealistic", "minimalist flat design" 3. Mention lighting: "golden hour lighting", "soft diffused light", "dramatic shadows" 4. Specify composition: "close-up", "wide angle", "from above", "centered"
Advanced Prompting Techniques
These techniques produce significantly better results for photorealistic and professional imagery:
Camera & Lens Specifications
Include specific photography parameters for authentic looks:
- Lens focal length: "85mm f/1.4 lens", "50mm prime", "35mm wide angle"
- Aperture: "shallow depth of field at f/1.8", "deep focus at f/8"
- Camera angle: "eye-level shot", "low angle looking up", "3/4 profile view"
Example: "Portrait shot with 85mm f/1.4 lens, shallow depth of field, subject sharp against soft bokeh background"
Lighting Architecture
Specify complete lighting setups for professional results:
- Three-point lighting: key light, fill light, rim/back light
- Catchlights: reflections in eyes for portraits
- Shadow quality: "soft shadows", "dramatic hard shadows", "subtle rim light"
Example: "Professional headshot with three-point lighting setup, soft key light from left, subtle fill light, rim light for hair separation, visible catchlights in eyes"
Film Stock & Era Aesthetics
Reference specific film stocks or eras for authentic period looks:
- Film stocks: "Kodak Portra 400 color palette", "Fujifilm Pro 400H tones", "Kodak Tri-X grain"
- Era aesthetics: "1990s disposable camera quality", "early-2000s digital camera look", "1970s Polaroid style"
- Grain and texture: "subtle film grain", "realistic sensor noise"
Example: "Casual portrait with Kodak Portra 400 color tones, natural film grain, 1990s aesthetic, soft warm highlights"
Facial Consistency (for portraits/people)
When generating or editing images with faces, explicitly state preservation requirements:
"Keep the facial features exactly consistent""Preserve original face structure and proportions""Do not alter or change the face"
Material & Texture Details
Add realistic texture specifications:
- Skin: "natural skin texture with visible pores", "subtle skin imperfections"
- Fabric: "fine wool texture visible", "silk sheen and drape"
- Surfaces: "brushed metal finish", "weathered wood grain"
Example: "Close-up portrait showing natural skin texture with visible pores, fine fabric detail on collar, realistic hair strands"
Composition Framing
Be precise about framing and subject positioning:
- Shot types: "chest-up framing", "3/4 body shot", "full body"
- Positioning: "centered subject", "rule of thirds placement", "mirror selfie angle"
- Aspect ratio context: "portrait orientation", "landscape format", "square crop"
Example Prompts by Use Case
UI/Frontend:
- "A modern dashboard UI mockup with dark theme, showing analytics charts"
- "Clean minimalist app icon for a task management app, rounded square shape"
- "Hero image for a SaaS landing page, abstract gradient with geometric shapes"
Documentation:
- "Simple architecture diagram showing microservices connected by arrows"
- "Flowchart illustrating user authentication process"
Placeholders:
- "Professional headshot placeholder, silhouette style, neutral gray background"
- "Product image placeholder, simple box shape with 'Image Coming Soon' text"
Marketing/Creative:
- "Isometric illustration of a modern office workspace"
- "Gradient abstract background suitable for presentation slides"
Professional Portraits (using advanced techniques):
- "Corporate headshot, 85mm f/2.8 lens, three-point studio lighting, navy blue suit, neutral gray backdrop, subtle catchlights, chest-up framing, natural skin texture"
- "Casual lifestyle portrait, Kodak Portra 400 tones, natural window light, soft shadows, 3/4 body shot, authentic film grain, early-2000s digital aesthetic"
E-commerce & Product Photography:
- "Product photo of leather watch, 100mm macro lens, soft diffused lighting, visible leather texture and stitching, clean white background, subtle reflection"
- "Fashion flat-lay, overhead shot, soft natural lighting, fabric texture visible, minimalist composition"
Troubleshooting
"GEMINI_API_KEY not set"
Ensure the environment variable is set in your current shell:
Windows (PowerShell):
echo $env:GEMINI_API_KEY # Should show your keymacOS/Linux:
echo $GEMINI_API_KEY # Should show your key"API request failed with HTTP status 400"
- Check your prompt for special characters that may break JSON
- Ensure the prompt isn't empty
- Verify API key is valid
"API request failed with HTTP status 429"
- Rate limited - wait a moment and retry
- Consider upgrading your API quota
"No image data found in response"
- The model may have refused the prompt (content policy)
- Try rephrasing the prompt
- Check if the model returned an error message in the response
Image is corrupted/won't open
- Ensure Python 3.6+ is installed
- Check if the full response was received (network issues)
- Verify output path is writable
Windows-specific issues
- Make sure Python is in your PATH
- Use forward slashes or escaped backslashes in paths
API Costs
Check Google AI pricing for current Gemini API costs. Image generation typically costs more than text generation.
Limitations
- Maximum prompt length varies by model
- Some content types may be restricted by Google's content policy
- Generated images are subject to Google's terms of service
- Rate limits apply based on your API tier
#!/usr/bin/env python3
"""
Imagen - Google Gemini Image Generation Script
Cross-platform image generation using Google Gemini API.
Works on Windows, macOS, and Linux.
Usage:
python generate_image.py "prompt" [output_path]
Environment variables:
GEMINI_API_KEY (required) - Your Google Gemini API key
IMAGE_SIZE (optional) - Image size: "512", "1K" (default), or "2K"
GEMINI_MODEL (optional) - Model ID (default: gemini-3-pro-image-preview)
"""
import argparse
import base64
import json
import os
import sys
import urllib.request
import urllib.error
from pathlib import Path
# Configuration
DEFAULT_MODEL_ID = "gemini-3-pro-image-preview"
API_BASE_URL = "https://generativelanguage.googleapis.com/v1beta/models"
DEFAULT_IMAGE_SIZE = "1K"
VALID_SIZES = {"512", "1K", "2K"}
def get_api_endpoint(model_id: str) -> str:
"""Build the API endpoint URL for the given model."""
# Use streamGenerateContent for image generation models
return f"{API_BASE_URL}/{model_id}:streamGenerateContent"
def get_api_key() -> str:
"""Get the Gemini API key from environment variable."""
api_key = os.environ.get("GEMINI_API_KEY")
if not api_key:
print("Error: GEMINI_API_KEY environment variable not set", file=sys.stderr)
print("\nTo set it:", file=sys.stderr)
print(" Windows (PowerShell): $env:GEMINI_API_KEY = 'your-key'", file=sys.stderr)
print(" Windows (CMD): set GEMINI_API_KEY=your-key", file=sys.stderr)
print(" macOS/Linux: export GEMINI_API_KEY='your-key'", file=sys.stderr)
print("\nGet a free key at: https://aistudio.google.com/", file=sys.stderr)
sys.exit(1)
return api_key
def validate_image_size(size: str) -> str:
"""Validate and return the image size."""
if size not in VALID_SIZES:
print(f"Warning: Invalid IMAGE_SIZE '{size}'. Using default '{DEFAULT_IMAGE_SIZE}'", file=sys.stderr)
return DEFAULT_IMAGE_SIZE
return size
def create_output_dir(output_path: Path) -> None:
"""Create output directory if it doesn't exist."""
output_dir = output_path.parent
if output_dir and not output_dir.exists():
output_dir.mkdir(parents=True, exist_ok=True)
def build_request_body(prompt: str, image_size: str) -> bytes:
"""Build the JSON request body for the API."""
request_data = {
"contents": [
{
"role": "user",
"parts": [
{"text": prompt}
]
}
],
"generationConfig": {
"responseModalities": ["IMAGE", "TEXT"],
"imageConfig": {
"image_size": image_size
}
}
}
return json.dumps(request_data).encode("utf-8")
def make_api_request(api_key: str, model_id: str, request_body: bytes) -> dict:
"""Make the API request and return the response."""
endpoint = get_api_endpoint(model_id)
url = f"{endpoint}?key={api_key}"
headers = {
"Content-Type": "application/json"
}
req = urllib.request.Request(url, data=request_body, headers=headers, method="POST")
try:
with urllib.request.urlopen(req, timeout=120) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as e:
error_body = e.read().decode("utf-8") if e.fp else ""
error_detail = ""
# Try to extract error message from response
if error_body:
try:
error_json = json.loads(error_body)
error_detail = error_json.get("error", {}).get("message", "")
except json.JSONDecodeError:
error_detail = error_body
# Provide user-friendly messages for common errors
if e.code == 429:
print("=" * 60, file=sys.stderr)
print("ERROR: Gemini API quota exhausted", file=sys.stderr)
print("=" * 60, file=sys.stderr)
print("\nYou've exceeded your Gemini API usage limits.", file=sys.stderr)
print("\nWhat to do:", file=sys.stderr)
print(" 1. Wait for your quota to reset (usually resets daily)", file=sys.stderr)
print(" 2. Check your usage at: https://aistudio.google.com/", file=sys.stderr)
print(" 3. Consider upgrading your API plan if needed", file=sys.stderr)
if error_detail:
print(f"\nAPI message: {error_detail}", file=sys.stderr)
elif e.code == 403:
print("=" * 60, file=sys.stderr)
print("ERROR: Gemini API access denied", file=sys.stderr)
print("=" * 60, file=sys.stderr)
print("\nYour API key is invalid or lacks required permissions.", file=sys.stderr)
print("\nWhat to do:", file=sys.stderr)
print(" 1. Verify your API key at: https://aistudio.google.com/", file=sys.stderr)
print(" 2. Ensure the Gemini API is enabled for your project", file=sys.stderr)
print(" 3. Check that your key has image generation permissions", file=sys.stderr)
if error_detail:
print(f"\nAPI message: {error_detail}", file=sys.stderr)
elif e.code == 400:
print("=" * 60, file=sys.stderr)
print("ERROR: Invalid request to Gemini API", file=sys.stderr)
print("=" * 60, file=sys.stderr)
print("\nThe request was rejected by the API.", file=sys.stderr)
print("\nPossible causes:", file=sys.stderr)
print(" - Prompt may contain blocked content", file=sys.stderr)
print(" - Prompt format may be invalid", file=sys.stderr)
print(" - Image generation may not be available for this prompt", file=sys.stderr)
if error_detail:
print(f"\nAPI message: {error_detail}", file=sys.stderr)
elif e.code >= 500:
print("=" * 60, file=sys.stderr)
print("ERROR: Gemini API server error", file=sys.stderr)
print("=" * 60, file=sys.stderr)
print(f"\nThe Gemini API returned a server error (HTTP {e.code}).", file=sys.stderr)
print("\nWhat to do:", file=sys.stderr)
print(" 1. Wait a few minutes and try again", file=sys.stderr)
print(" 2. Check Gemini API status if the issue persists", file=sys.stderr)
if error_detail:
print(f"\nAPI message: {error_detail}", file=sys.stderr)
else:
print(f"Error: API request failed with HTTP status {e.code}", file=sys.stderr)
if error_detail:
print(f"API message: {error_detail}", file=sys.stderr)
elif error_body:
print(f"Response: {error_body}", file=sys.stderr)
sys.exit(1)
except urllib.error.URLError as e:
print("=" * 60, file=sys.stderr)
print("ERROR: Failed to connect to Gemini API", file=sys.stderr)
print("=" * 60, file=sys.stderr)
print(f"\nConnection error: {e.reason}", file=sys.stderr)
print("\nWhat to do:", file=sys.stderr)
print(" 1. Check your internet connection", file=sys.stderr)
print(" 2. Verify the API endpoint is accessible", file=sys.stderr)
print(" 3. Check if a firewall or proxy is blocking the request", file=sys.stderr)
sys.exit(1)
def extract_image_data(response: dict) -> str:
"""Extract base64 image data from the API response."""
try:
# Handle both streaming array and single object responses
if isinstance(response, list):
candidates = response[0].get("candidates", [])
else:
candidates = response.get("candidates", [])
if not candidates:
raise ValueError("No candidates in response")
parts = candidates[0].get("content", {}).get("parts", [])
for part in parts:
if "inlineData" in part:
return part["inlineData"].get("data", "")
raise ValueError("No image data found in response parts")
except (KeyError, IndexError, TypeError) as e:
print(f"Error: Failed to parse response: {e}", file=sys.stderr)
print(f"Response: {json.dumps(response, indent=2)}", file=sys.stderr)
sys.exit(1)
def save_image(image_data: str, output_path: Path) -> None:
"""Decode and save the base64 image data."""
try:
image_bytes = base64.b64decode(image_data)
output_path.write_bytes(image_bytes)
except Exception as e:
print(f"Error: Failed to save image: {e}", file=sys.stderr)
sys.exit(1)
def get_file_size(path: Path) -> str:
"""Get human-readable file size."""
size = path.stat().st_size
for unit in ["B", "KB", "MB", "GB"]:
if size < 1024:
return f"{size:.1f} {unit}"
size /= 1024
return f"{size:.1f} TB"
def main():
parser = argparse.ArgumentParser(
description="Generate images using Google Gemini AI",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
python generate_image.py "A sunset over mountains"
python generate_image.py "An app icon" ./icons/app.png
python generate_image.py --size 2K "High-res landscape" ./wallpaper.png
python generate_image.py --model gemini-3-pro-image-preview "A logo" ./logo.png
Environment Variables:
GEMINI_API_KEY Your Google Gemini API key (required)
IMAGE_SIZE Image size: 512, 1K (default), or 2K
GEMINI_MODEL Model ID for image generation
"""
)
parser.add_argument("prompt", help="Text description of the image to generate")
parser.add_argument("output", nargs="?", default="./generated-image.png",
help="Output file path (default: ./generated-image.png)")
parser.add_argument("--size", choices=["512", "1K", "2K"],
help="Image size (overrides IMAGE_SIZE env var)")
parser.add_argument("--model", "-m",
help=f"Gemini model ID (default: {DEFAULT_MODEL_ID})")
args = parser.parse_args()
# Get configuration
api_key = get_api_key()
model_id = args.model or os.environ.get("GEMINI_MODEL", DEFAULT_MODEL_ID)
image_size = args.size or os.environ.get("IMAGE_SIZE", DEFAULT_IMAGE_SIZE)
image_size = validate_image_size(image_size)
output_path = Path(args.output)
# Create output directory
create_output_dir(output_path)
# Display info
print(f"Generating image with prompt: \"{args.prompt}\"")
print(f"Model: {model_id}")
print(f"Image size: {image_size}")
print(f"Output path: {output_path}")
print()
# Build and send request
request_body = build_request_body(args.prompt, image_size)
response = make_api_request(api_key, model_id, request_body)
# Extract and save image
image_data = extract_image_data(response)
if not image_data:
print("Error: No image data received from API", file=sys.stderr)
sys.exit(1)
save_image(image_data, output_path)
# Verify and report success
if output_path.exists() and output_path.stat().st_size > 0:
file_size = get_file_size(output_path)
print("Success! Image generated and saved.")
print(f"File: {output_path}")
print(f"Size: {file_size}")
else:
print(f"Error: Failed to save image to {output_path}", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()
Related skills
Forks & variants (1)
Imagen has 1 known copy in the catalog totaling 29 installs. They canonicalize to this original listing.
- sanjay3290 - 29 installs
How it compares
Use imagen for quick Gemini-generated session visuals instead of exporting from design tools when fidelity requirements are exploratory.
FAQ
Which model does imagen use?
imagen uses Google Gemini's `gemini-3-pro-image-preview` image generation model. The skill is version 1.0 under Apache-2.0 from sanjay3290/ai-skills for Claude Code sessions.
What can imagen generate?
imagen creates UI mockups, icons, illustrations, diagrams, concept art, and placeholder images on demand. Invoke it whenever a Claude Code session needs visual assets alongside frontend or documentation work.