
Gemini Video Understanding
- 35 installs
- 1 repo stars
- Updated November 15, 2025
- aia-11-hn-mib/mib-mockinterviewaibot
gemini-video-understanding is a Claude Code skill that analyzes videos with Google's Gemini API to summarize, transcribe, and answer questions about video content.
About
gemini-video-understanding is a Claude Code skill for analyzing video with Google's Gemini API. A developer uses it to summarize videos, answer questions about content, transcribe audio with visual descriptions and timestamps, or process YouTube URLs. It ships an analyze_video.py script, supports clipping by start/end offset, and can compare up to 10 videos on Gemini 2.5.
- Analyzes videos with Gemini: summarize, answer questions, transcribe with timestamps
- Supports 9 video formats, YouTube URLs, video clipping, and custom frame rates
- Ships an analyze_video.py script and handles context windows up to 2M tokens (~6 hours)
Gemini Video Understanding by the numbers
- 35 all-time installs (skills.sh)
- Ranked #948 of 1,337 Generative Media skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
gemini-video-understanding capabilities & compatibility
Requires a Gemini API key; token cost scales with video length and resolution.
- Capabilities
- video analysis · video summarization · transcription · video qa
- Works with
- gcp
- Use cases
- video generation · transcription · data analysis
- Pricing
- Bring your own API key
What gemini-video-understanding says it does
This skill enables comprehensive video analysis using Google's Gemini API, including video summarization, question answering, transcription, timestamp references, and more.
MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
npx skills add https://github.com/aia-11-hn-mib/mib-mockinterviewaibot --skill gemini-video-understandingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 35 |
|---|---|
| repo stars | ★ 1 |
| Last updated | November 15, 2025 |
| Repository | aia-11-hn-mib/mib-mockinterviewaibot ↗ |
What it does
Summarize, transcribe, and answer questions about video content using the Gemini API.
Who is it for?
Summarizing videos, transcribing with timestamps, answering questions about content, and processing YouTube URLs.
Skip if: Comparing many videos at once on older models, since multi-video comparison requires Gemini 2.5 or newer.
When should I use this skill?
Analyzing, summarizing, transcribing, or answering questions about video content, including YouTube.
What you get
Videos and YouTube URLs are turned into summaries, timestamped transcripts, and answers.
By the numbers
- 9 supported video formats
- up to 10 videos compared on Gemini 2.5+
- context windows up to 2M tokens (~6 hours)
Files
Gemini Video Understanding Skill
This skill enables comprehensive video analysis using Google's Gemini API, including video summarization, question answering, transcription, timestamp references, and more.
Capabilities
- Video Summarization: Create concise summaries of video content
- Question Answering: Answer specific questions about video content
- Transcription: Transcribe audio with visual descriptions and timestamps
- Timestamp References: Query specific moments in videos (MM:SS format)
- Video Clipping: Process specific segments using start/end offsets
- Multiple Videos: Compare and analyze up to 10 videos (Gemini 2.5+)
- YouTube Support: Analyze YouTube videos directly (preview feature)
- Custom Frame Rate: Adjust FPS sampling for different video types
Supported Formats
- MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
Models Available
Gemini 2.5 Series:
gemini-2.5-pro- Best quality, 1M contextgemini-2.5-flash- Balanced quality/speed, 1M contextgemini-2.5-flash-preview-09-2025- Preview features, 1M context
Gemini 2.0 Series:
gemini-2.0-flash- Fast processinggemini-2.0-flash-lite- Lightweight option
Context Windows:
- 2M token models: ~2 hours (default) or ~6 hours (low-res)
- 1M token models: ~1 hour (default) or ~3 hours (low-res)
API Key Configuration
The skill supports both Google AI Studio and Vertex AI endpoints.
Option 1: Google AI Studio (Default)
The skill checks for GEMINI_API_KEY in this order: 1. Process environment: process.env.GEMINI_API_KEY or $GEMINI_API_KEY 2. Project root: .env 3. .claude directory: .claude/.env 4. .claude/skills directory: .claude/skills/.env 5. Skill directory: .claude/skills/gemini-video-understanding/.env
Get your API key: https://aistudio.google.com/apikey
To set up:
# Environment variable (recommended)
export GEMINI_API_KEY="your-api-key-here"
# Or in .env file
echo "GEMINI_API_KEY=your-api-key-here" > .envOption 2: Vertex AI
To use Vertex AI instead:
# Enable Vertex AI
export GEMINI_USE_VERTEX=true
export VERTEX_PROJECT_ID=your-gcp-project-id
export VERTEX_LOCATION=us-central1 # Optional, defaults to us-central1Or in .env file:
GEMINI_USE_VERTEX=true
VERTEX_PROJECT_ID=your-gcp-project-id
VERTEX_LOCATION=us-central1Usage Instructions
When to Use This Skill
Use this skill when the user asks to:
- Analyze, summarize, or describe video content
- Answer questions about videos
- Transcribe video audio with visual context
- Extract information from specific timestamps
- Compare multiple videos
- Process YouTube video content
- Create quizzes or educational content from videos
Basic Video Analysis
For video files:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Summarize this video in 3 key points"For YouTube URLs:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--youtube-url "https://www.youtube.com/watch?v=VIDEO_ID" \
--prompt "What are the main topics discussed?"Advanced Features
Video Clipping (specific time range):
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Summarize this segment" \
--start-offset "40s" \
--end-offset "80s"Custom Frame Rate:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Analyze the rapid movements" \
--fps 5Transcription with Timestamps:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Transcribe the audio with timestamps and visual descriptions"Multiple Videos (Gemini 2.5+ only):
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-paths "/path/video1.mp4" "/path/video2.mp4" \
--prompt "Compare these two videos and highlight the differences"Model Selection:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Detailed analysis" \
--model "gemini-2.5-pro"Script Parameters
Required (one of):
--video-path PATH Path to local video file
--youtube-url URL YouTube video URL
--video-paths PATH [PATH..] Multiple video paths (Gemini 2.5+)
Required:
--prompt TEXT Analysis prompt/question
Optional:
--model NAME Model to use (default: gemini-2.5-flash)
--start-offset TIME Video clip start (e.g., "40s", "1m30s")
--end-offset TIME Video clip end (e.g., "80s", "2m")
--fps NUMBER Frame sampling rate (default: 1)
--output-file PATH Save response to file
--verbose Show detailed processing infoCommon Use Cases
1. Video Summarization
Prompt: "Summarize this video in 3 key points with timestamps"2. Educational Content
Prompt: "Create a quiz with 5 questions and answer key based on this video"3. Timestamp-Specific Questions
Prompt: "What happens at 01:15 and how does it relate to the topic at 02:30?"4. Transcription
Prompt: "Transcribe the audio from this video with timestamps for salient events and visual descriptions"5. Content Comparison
Prompt: "Compare these two product demo videos. Which one explains the features more clearly?"6. Action Detection
Prompt: "List all the actions performed in this tutorial video with timestamps"Rate Limits & Quotas
Free Tier (per model):
- 10-15 RPM (requests per minute)
- 1M-4M TPM (tokens per minute)
- 1,500 RPD (requests per day)
YouTube Limitations:
- Free tier: 8 hours of YouTube video per day
- Paid tier: No length-based limits
- Public videos only (no private/unlisted)
Storage (Files API):
- 20GB per project
- 2GB per file
- 48-hour retention period
Token Calculation
Video tokens depend on resolution:
- Default resolution: ~300 tokens per second of video
- Low resolution: ~100 tokens per second of video
Example: A 10-minute video = 600 seconds × 300 tokens = ~180,000 tokens
Error Handling
Common errors and solutions:
| Error | Cause | Solution |
|---|---|---|
| 400 Bad Request | Invalid video format or corrupt file | Check file format and integrity |
| 403 Forbidden | Invalid/missing API key | Verify GEMINI_API_KEY configuration |
| 404 Not Found | File URI not found | Ensure file is uploaded and active |
| 429 Too Many Requests | Rate limit exceeded | Implement backoff, upgrade to paid tier |
| 500 Internal Error | Server-side issue | Retry with exponential backoff |
Best Practices
1. Use Files API for videos >20MB - More reliable than inline data 2. Wait for file processing - Poll until state is ACTIVE before analysis 3. Optimize FPS - Use lower FPS for static content to save tokens 4. Clip long videos - Process specific segments instead of entire video 5. Cache context - Reuse uploaded files for multiple queries 6. Batch processing - Process multiple short videos in one request (2.5+) 7. Specific prompts - Be precise about what you want to extract
Implementation Notes
For Claude Code:
When a user requests video analysis:
1. Check API key availability first using the helper script 2. Determine video source: local file, YouTube URL, or multiple videos 3. Select appropriate model based on requirements (default: gemini-2.5-flash) 4. Run the analysis script with proper parameters 5. Parse and present results to the user clearly 6. Handle errors gracefully with helpful suggestions
Files API Workflow:
For videos >20MB or reusable content: 1. Upload video using Files API (script handles this automatically) 2. Wait for ACTIVE state (polling included in script) 3. Use file URI for analysis 4. Files auto-delete after 48 hours
Inline Data Workflow:
For videos <20MB: 1. Read video file as bytes 2. Base64 encode for API 3. Send in generateContent request 4. Single-use, no upload needed
Example Workflows
Workflow 1: YouTube Video Summary
# User: "Analyze this YouTube tutorial video"
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--youtube-url "https://www.youtube.com/watch?v=abc123" \
--prompt "Create a structured summary with: 1) Main topics, 2) Key takeaways, 3) Recommended audience"Workflow 2: Interview Transcription
# User: "Transcribe this interview with timestamps"
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "interview.mp4" \
--prompt "Transcribe this interview with speaker labels, timestamps, and visual descriptions of gestures or slides shown"Workflow 3: Product Comparison
# User: "Compare these two product demo videos"
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-paths "demo1.mp4" "demo2.mp4" \
--model "gemini-2.5-pro" \
--prompt "Compare these product demos on: features shown, presentation quality, clarity of explanation, and overall effectiveness"Troubleshooting
API Key Not Found:
# Check API key detection
python .claude/skills/gemini-video-understanding/scripts/check_api_key.pyVideo Too Large:
Error: Request size exceeds 20MB
Solution: Script automatically uses Files API for large videosProcessing Timeout:
Error: File not reaching ACTIVE state
Solution: Check video integrity, try smaller file, or different formatRate Limit Errors:
Error: 429 Too Many Requests
Solution: Wait before retry, or upgrade to paid tierAdditional Resources
- API Documentation: https://ai.google.dev/gemini-api/docs/video-understanding
- Files API Guide: https://ai.google.dev/gemini-api/docs/vision#uploading-files
- Rate Limits: https://ai.google.dev/gemini-api/docs/rate-limits
- Pricing: https://ai.google.dev/pricing
- Get API Key: https://aistudio.google.com/apikey
Version History
- 1.0.0 (2025-10-26): Initial release with full video understanding capabilities
# Gemini Video Understanding - API Configuration
# Copy this file to .env and configure your API settings
# ==== Google AI Studio (Default) ====
# Get your API key from: https://aistudio.google.com/apikey
GEMINI_API_KEY=your_api_key_here
# ==== Vertex AI (Optional) ====
# Uncomment to use Vertex AI instead of AI Studio
# GEMINI_USE_VERTEX=true
# VERTEX_PROJECT_ID=your-gcp-project-id
# VERTEX_LOCATION=us-central1
Gemini Video Understanding - Examples
Setup
1. Install dependencies:
pip install -r .claude/skills/gemini-video-understanding/requirements.txt2. Configure API key:
# Option 1: Environment variable (recommended)
export GEMINI_API_KEY="your-api-key-here"
# Option 2: Copy and edit .env file
cp .claude/skills/gemini-video-understanding/.env.example \
.claude/skills/gemini-video-understanding/.env
# Then edit the .env file with your API key3. Verify configuration:
python .claude/skills/gemini-video-understanding/scripts/check_api_key.pyBasic Examples
Example 1: Simple Video Summary
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Summarize this video in 3 key points"Output:
1. The video demonstrates how to use the Gemini API for video analysis
2. It covers three main input methods: Files API, inline data, and YouTube URLs
3. Examples show various use cases including transcription and timestamp queriesExample 2: YouTube Video Analysis
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--youtube-url "https://www.youtube.com/watch?v=dQw4w9WgXcQ" \
--prompt "What is this video about? List the main topics discussed."Example 3: Detailed Transcription
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/interview.mp4" \
--prompt "Transcribe this interview with timestamps and speaker labels" \
--verboseAdvanced Examples
Example 4: Video Clipping
Analyze only a specific portion of a video:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/long-video.mp4" \
--prompt "Summarize this segment" \
--start-offset "2m30s" \
--end-offset "5m15s"Example 5: High Frame Rate Analysis
For videos with fast action:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/sports.mp4" \
--prompt "Describe the key movements and techniques shown" \
--fps 5Example 6: Comparing Multiple Videos
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-paths "/path/video1.mp4" "/path/video2.mp4" \
--prompt "Compare these two product demos. Which one is more effective and why?" \
--model "gemini-2.5-pro"Example 7: Educational Content Generation
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--youtube-url "https://www.youtube.com/watch?v=VIDEO_ID" \
--prompt "Create a quiz with 5 multiple-choice questions based on this video. Include an answer key with explanations."Example 8: Timestamp-Specific Questions
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/tutorial.mp4" \
--prompt "What happens at 01:15 and 02:30? How are these two moments related?"Prompting Best Practices
For Summaries
"Provide a structured summary with:
1. Main topic or theme
2. Key points discussed (3-5 bullet points)
3. Notable quotes or moments with timestamps
4. Target audience or intended purpose"For Transcriptions
"Transcribe this video including:
- Speaker labels (Speaker 1, Speaker 2, etc.)
- Timestamps for each segment (MM:SS format)
- Visual descriptions in [brackets] when relevant
- Key gestures or on-screen text mentioned"For Educational Content
"Analyze this educational video and create:
1. Learning objectives (3-5 points)
2. Key concepts explained with timestamps
3. 5 quiz questions with multiple choice answers
4. Answer key with brief explanations"For Comparisons
"Compare these videos on the following criteria:
1. Content quality and accuracy
2. Presentation style and clarity
3. Production quality (audio, video, editing)
4. Target audience suitability
5. Overall effectiveness
Provide ratings (1-5) for each criterion with justification."Use Case Examples
Use Case 1: Meeting Notes
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "team-meeting.mp4" \
--prompt "Create meeting notes with: 1) Attendees (if visible), 2) Topics discussed with timestamps, 3) Action items identified, 4) Decisions made" \
--output-file "meeting-notes.txt"Use Case 2: Content Moderation
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "user-submitted.mp4" \
--prompt "Analyze this video for: 1) Overall content type, 2) Any inappropriate content, 3) Compliance with community guidelines, 4) Recommended action (approve/review/reject)"Use Case 3: Accessibility - Video Descriptions
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "promotional-video.mp4" \
--prompt "Create an audio description track for visually impaired viewers. Include descriptions of: visual elements, on-screen text, scene changes, and important actions at appropriate timestamps."Use Case 4: Sports Analysis
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "game-footage.mp4" \
--prompt "Analyze this game footage: 1) Key plays with timestamps, 2) Player performance highlights, 3) Strategic decisions, 4) Turning points in the game" \
--fps 2 \
--verboseUse Case 5: Tutorial Enhancement
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--youtube-url "https://www.youtube.com/watch?v=TUTORIAL_ID" \
--prompt "Create an enhanced tutorial outline with: 1) Chapter markers with timestamps, 2) Prerequisites mentioned, 3) Tools/resources needed, 4) Common mistakes to avoid, 5) Practice exercises suggested"Output Options
Save to File
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "video.mp4" \
--prompt "Summarize this video" \
--output-file "summary.txt"JSON Output
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "video.mp4" \
--prompt "Summarize this video" \
--json \
--output-file "summary.json"Verbose Mode (with token usage)
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "video.mp4" \
--prompt "Analyze this video" \
--verboseError Handling
Check API Key
python .claude/skills/gemini-video-understanding/scripts/check_api_key.pyCommon Errors
File Not Found:
# Error: Video file not found: /path/to/video.mp4
# Solution: Check file path and ensure file exists
ls -lh /path/to/video.mp4API Key Invalid:
# Error: 403 Forbidden
# Solution: Verify API key is correct
python .claude/skills/gemini-video-understanding/scripts/check_api_key.pyRate Limit:
# Error: 429 Too Many Requests
# Solution: Wait before retrying, or upgrade to paid tierModel Selection
Fast Processing
--model "gemini-2.0-flash-lite" # Fastest, good for simple tasksBalanced (Default)
--model "gemini-2.5-flash" # Best balance of speed and qualityHighest Quality
--model "gemini-2.5-pro" # Most accurate, slower, higher token usagePerformance Tips
1. Use video clipping to analyze only relevant portions 2. Adjust FPS based on content type (lower for static, higher for action) 3. Use Files API for videos >20MB (handled automatically) 4. Batch process multiple short videos in one request (Gemini 2.5+) 5. Cache uploaded files - reuse file URIs for multiple queries
Resources
- API Documentation: https://ai.google.dev/gemini-api/docs/video-understanding
- Get API Key: https://aistudio.google.com/apikey
- Pricing: https://ai.google.dev/pricing
- Rate Limits: https://ai.google.dev/gemini-api/docs/rate-limits
Quick Start Guide - Gemini Video Understanding
1. Install Dependencies
pip install google-genaiOr use the requirements file:
pip install -r .claude/skills/gemini-video-understanding/requirements.txt2. Configure API Key
Choose one option:
Option A: Environment Variable (Recommended)
export GEMINI_API_KEY="your-api-key-here"Option B: Skill Directory .env
cp .claude/skills/gemini-video-understanding/.env.example \
.claude/skills/gemini-video-understanding/.env
# Edit the .env file with your API keyOption C: Project Root .env
echo "GEMINI_API_KEY=your-api-key-here" > .envGet your API key at: https://aistudio.google.com/apikey
3. Verify Setup
python .claude/skills/gemini-video-understanding/scripts/check_api_key.py4. Analyze Your First Video
Local video file:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "/path/to/video.mp4" \
--prompt "Summarize this video in 3 sentences"YouTube video:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--youtube-url "https://www.youtube.com/watch?v=VIDEO_ID" \
--prompt "What are the main topics discussed?"5. Common Commands
Transcribe with timestamps:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "video.mp4" \
--prompt "Transcribe this video with timestamps and visual descriptions"Analyze specific segment:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "video.mp4" \
--prompt "Summarize this part" \
--start-offset "1m30s" \
--end-offset "3m"Compare videos:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-paths "video1.mp4" "video2.mp4" \
--prompt "Compare these videos"Save to file:
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
--video-path "video.mp4" \
--prompt "Analyze this video" \
--output-file "analysis.txt"Help
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py --helpSee EXAMPLES.md for more detailed examples and use cases.
Gemini Video Understanding Skill
A comprehensive Claude Code skill for analyzing videos using Google's Gemini API.
Overview
This skill enables AI-powered video analysis with capabilities including:
- Video summarization and content description
- Question answering about video content
- Audio transcription with visual descriptions
- Timestamp-based queries (MM:SS format)
- Video clipping (specific time ranges)
- Multiple video comparison (Gemini 2.5+)
- YouTube video processing
- Custom frame rate sampling
Features
✅ Three Video Input Methods:
- Local files (Files API for >20MB, inline for <20MB)
- YouTube URLs (preview feature)
- Multiple videos (up to 10 with Gemini 2.5+)
✅ Advanced Processing:
- Video clipping with start/end offsets
- Custom FPS sampling (default: 1 FPS)
- Context windows up to 2M tokens (~6 hours of video)
✅ 9 Supported Video Formats: MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
✅ Multiple Models:
- Gemini 2.5 Series (Pro, Flash, Preview)
- Gemini 2.0 Series (Flash, Flash-Lite)
✅ Flexible API Key Management: Checks in order: Process env → Skill directory → Project root
Quick Start
1. Install: pip install google-genai 2. Configure: export GEMINI_API_KEY="your-key" 3. Verify: python scripts/check_api_key.py 4. Analyze: python scripts/analyze_video.py --video-path video.mp4 --prompt "Summarize"
See QUICKSTART.md for detailed setup instructions.
Documentation
- [SKILL.md](SKILL.md) - Complete skill documentation
- [QUICKSTART.md](QUICKSTART.md) - Quick setup guide
- [EXAMPLES.md](EXAMPLES.md) - Comprehensive examples and use cases
Directory Structure
gemini-video-understanding/
├── SKILL.md # Main skill documentation
├── README.md # This file
├── QUICKSTART.md # Quick start guide
├── EXAMPLES.md # Detailed examples
├── requirements.txt # Python dependencies
├── .env.example # API key template
└── scripts/
├── analyze_video.py # Main video analysis script
└── check_api_key.py # API key verification toolCommon Use Cases
| Use Case | Command |
|---|---|
| Video Summary | --prompt "Summarize this video in 3 points" |
| Transcription | --prompt "Transcribe with timestamps and visual descriptions" |
| YouTube Analysis | --youtube-url "URL" --prompt "What are the main topics?" |
| Time Range | --start-offset "1m30s" --end-offset "3m" |
| Multi-Video | --video-paths vid1.mp4 vid2.mp4 --prompt "Compare" |
| High Quality | --model "gemini-2.5-pro" |
API Key Configuration
The skill automatically checks for GEMINI_API_KEY in this order:
1. Process Environment - $GEMINI_API_KEY 2. Skill Directory - .claude/skills/gemini-video-understanding/.env 3. Project Root - .env file
Get your API key at: https://aistudio.google.com/apikey
Rate Limits & Pricing
Free Tier:
- 10-15 requests per minute
- 1-4M tokens per minute
- 1,500 requests per day
- YouTube: 8 hours per day
Token Usage:
- Default: ~300 tokens/second of video
- Low-res: ~100 tokens/second of video
See full pricing: https://ai.google.dev/pricing
Support
- API Documentation: https://ai.google.dev/gemini-api/docs/video-understanding
- Files API: https://ai.google.dev/gemini-api/docs/vision#uploading-files
- Rate Limits: https://ai.google.dev/gemini-api/docs/rate-limits
- Issue Tracker: [Report issues on GitHub]
License
MIT License - See skill metadata for details.
Version
1.0.0 (2025-10-26)
google-genai>=0.3.0
Related skills
FAQ
Which video formats are supported?
The docs list MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, and 3GPP.
How long a video can it handle?
Docs state 2M token models cover about 2 hours default or 6 hours low-res, and 1M token models about 1 hour.
Can it process YouTube links?
Yes, YouTube URLs are supported as a preview feature.