Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
duc01226 avatar

Ai Multimodal

  • 34 installs
  • 7 repo stars
  • Updated June 18, 2026
  • duc01226/easyplatform

Processes vision, audio, video, document, and image content with Gemini multimodal APIs, including image and video generation.

About

Processes multimedia content using Gemini vision, audio, video, document, image, and generation APIs. A developer uses it when a task needs multimodal AI over media inputs.

  • Processes vision, audio, video, document, and image content via Gemini APIs
  • Includes image and video-generation capabilities

Ai Multimodal by the numbers

  • 34 all-time installs (skills.sh)
  • Ranked #8,855 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Jul 29, 2026 (Skillselion catalog sync)
npx skills add https://github.com/duc01226/easyplatform --skill ai-multimodal

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs34
repo stars7
Last updatedJune 18, 2026
Repositoryduc01226/easyplatform

What it does

Processes vision, audio, video, document, and image content with Gemini multimodal APIs, including image and video generation.

Files

SKILL.mdMarkdownGitHub ↗
Codex compatibility note:

>

- Invoke repository skills with $skill-name in Codex; this mirrored copy rewrites legacy Claude /skill-name references.
- Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
- User-question prompts mean to ask the user directly in Codex.
- Ignore Claude-specific mode-switch instructions when they appear.
- Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
- Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required spawn_agent subagent(s) for that task.
- Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
- For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
- If a required step/tool cannot run in this environment, stop and ask the user before adapting.

<!-- CODEX:PROJECT-REFERENCE-LOADING:START -->

Codex Project-Reference Loading (No Hooks)

Codex does not receive Claude hook-based doc injection. When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.

Always read:

  • docs/project-config.json (project-specific paths, commands, modules, and workflow/test settings)
  • docs/project-reference/docs-index-reference.md (routes to the full docs/project-reference/* catalog)
  • docs/project-reference/lessons.md (always-on guardrails and anti-patterns)

Situation-based docs:

  • Backend/CQRS/API/domain/entity changes: backend-patterns-reference.md, domain-entities-reference.md, project-structure-reference.md
  • Frontend/UI/styling/design-system: frontend-patterns-reference.md, scss-styling-guide.md, design-system/README.md
  • Spec/test-case planning or TC mapping: feature-docs-reference.md
  • Integration test implementation/review: integration-test-reference.md
  • E2E test implementation/review: e2e-test-reference.md
  • Code review/audit work: code-review-rules.md plus domain docs above based on changed files

Do not read all docs blindly. Start from docs-index-reference.md, then open only relevant files for the task.

<!-- CODEX:PROJECT-REFERENCE-LOADING:END -->

Quick Summary

Goal: Process and generate multimedia content (images, audio, video, documents) using Google Gemini API via Python scripts.

Workflow:

1. Identify Modality — Match input type to task (analyze, transcribe, extract, generate) 2. Check Limits — Inline max 20MB, File API max 2GB; split large audio at 15min chunks 3. Execute — Run gemini_batch_process.py with appropriate task and files 4. Post-Process — Format output as markdown with timestamps, save generated content

Key Rules:

  • Requires GEMINI_API_KEY environment variable
  • Always request specific nodes/files, avoid full-file downloads
  • Use media_optimizer.py to compress/split files exceeding limits

Be skeptical. Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence percentages (Idea should be more than 80%).

AI Multimodal

Purpose

Process audio, images, videos, and documents or generate images/videos using Google Gemini's multimodal API via bundled Python scripts.

When to Use

  • Analyzing images or screenshots (Gemini vision is preferred over Claude's built-in vision for complex tasks)
  • Transcribing audio files (meetings, podcasts, interviews)
  • Extracting data from PDFs, scanned documents, or charts
  • Processing video content (scene detection, temporal Q&A)
  • Generating images with Imagen 4 or videos with Veo 3
  • Converting documents to markdown with visual understanding

When NOT to Use

  • Simple text-only LLM calls -- use Claude directly
  • Reading a file Claude can already read (code, markdown, JSON) -- use Read tool
  • Building AI-powered application features -- use api-design or frontend-design
  • Music composition workflows -- load references/music-generation.md only when specifically requested
  • General prompt engineering -- use ai-artist skill

Prerequisites

export GEMINI_API_KEY="your-key"  # From https://aistudio.google.com/apikey
pip install google-genai python-dotenv pillow
python scripts/check_setup.py  # Verify setup

Optional: API key rotation for rate limits (set GEMINI_API_KEY_2, GEMINI_API_KEY_3).

Workflow

Step 1: Identify Modality

Input TypeTaskCommand
Image (PNG/JPG/WEBP)Analyze, caption, OCR--task analyze
Audio (WAV/MP3/AAC)Transcribe, summarize--task transcribe
Video (MP4/MOV)Scene detection, Q&A--task analyze
PDF/DocumentExtract tables, forms--task extract
Text promptGenerate image--task generate
Text promptGenerate video--task generate-video

Step 2: Check Limits

  • Inline upload: max 20MB
  • File API: max 2GB (auto-used for large files)
  • Audio transcription: split at 15-minute chunks for full transcript
  • Video transcription: extract audio first, then split and transcribe
  • Formats: Audio (WAV/MP3/AAC, up to 9.5h), Images (PNG/JPEG/WEBP, up to 3.6k), Video (MP4/MOV, up to 6h), PDF (up to 1k pages)

IF file exceeds limits, use scripts/media_optimizer.py to compress/split first.

Step 3: Execute

Quick check: If gemini CLI is available, use: "<prompt>" | gemini -y -m gemini-2.5-flash

Standard: Use the batch processing script:

# Analyze media
python scripts/gemini_batch_process.py --files <file> --task <analyze|transcribe|extract>

# Generate content
python scripts/gemini_batch_process.py --task generate --prompt "description"
python scripts/gemini_batch_process.py --task generate-video --prompt "description"

Stdin support: cat image.png | python scripts/gemini_batch_process.py --task analyze --prompt "Describe this"

Step 4: Post-Processing

  • For transcripts: output in markdown with [HH:MM:SS -> HH:MM:SS] timestamps
  • For document extraction: save as structured markdown under docs/assets/
  • For generated images/videos: save to working directory with descriptive filename

Step 5: Verification

  • Confirm output matches expected format and completeness
  • For long transcripts: verify no truncation occurred (check chunk boundaries)
  • For generated content: verify quality meets prompt requirements

Models

PurposeModelNotes
Analysis (fast)gemini-2.5-flashRecommended default
Analysis (advanced)gemini-2.5-proComplex reasoning tasks
Image generationimagen-4.0-generate-001Standard quality
Image generation (quality)imagen-4.0-ultra-generate-001Best quality
Image generation (speed)imagen-4.0-fast-generate-001Fastest
Video generationveo-3.1-generate-preview8s clips with audio

Scripts Reference

  • `gemini_batch_process.py` -- CLI orchestrator for all tasks, auto-resolves API keys and models
  • `media_optimizer.py` -- Compress/resize/split media to fit Gemini limits
  • `document_converter.py` -- Convert PDFs/images/Office docs to markdown
  • `check_setup.py` -- Verify environment, dependencies, and API key

Use --help on any script for full options.

Examples

Example 1: Transcribe a Meeting Recording

Input: 45-minute meeting audio file meeting-2025-01-15.mp3

Steps:

1. File is >15min, so split first:

    python scripts/media_optimizer.py --input meeting-2025-01-15.mp3 --split-duration 900

2. Transcribe each chunk:

    python scripts/gemini_batch_process.py --files meeting-part-*.mp3 --task transcribe

3. Output: Markdown file with timestamps, speaker detection, and metadata (duration, topics covered)

Example 2: Extract Data from a PDF Report

Input: Quarterly HR report PDF with tables, charts, and forms

Steps:

1. Convert and extract:

    python scripts/document_converter.py --input quarterly-report.pdf --output docs/assets/

2. Output: Structured markdown with tables preserved, chart descriptions, and form field values extracted

Detailed References

Load for in-depth guidance:

TopicFile
Audio processingreferences/audio-processing.md
Vision/image analysisreferences/vision-understanding.md
Image generationreferences/image-generation.md
Video analysisreferences/video-analysis.md
Video generationreferences/video-generation.md
Music generationreferences/music-generation.md

Related Skills

  • ai-artist -- for prompt engineering and optimization (not media processing)
  • media-processing -- for FFmpeg-based audio/video encoding without AI
  • pdf-to-markdown -- for simple PDF text extraction without vision AI

---

[IMPORTANT] Use task tracking to break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ATTENTION ask user whether to skip.

<!-- SYNC:ai-mistake-prevention -->

AI Mistake Prevention — Failure modes to avoid on every task:

>

Check downstream references before deleting. Deleting components causes documentation and code staleness cascades. Map all referencing files before removal.
Verify AI-generated content against actual code. AI hallucinates APIs, class names, and method signatures. Always grep to confirm existence before documenting or referencing.
Trace full dependency chain after edits. Changing a definition misses downstream variables and consumers derived from it. Always trace the full chain.
Trace ALL code paths when verifying correctness. Confirming code exists is not confirming it executes. Always trace early exits, error branches, and conditional skips — not just happy path.
When debugging, ask "whose responsibility?" before fixing. Trace whether bug is in caller (wrong data) or callee (wrong handling). Fix at responsible layer — never patch symptom site.
Assume existing values are intentional — ask WHY before changing. Before changing any constant, limit, flag, or pattern: read comments, check git blame, examine surrounding code.
Verify ALL affected outputs, not just the first. Changes touching multiple stacks require verifying EVERY output. One green check is not all green checks.
Holistic-first debugging — resist nearest-attention trap. When investigating any failure, list EVERY precondition first (config, env vars, DB names, endpoints, DI registrations, data preconditions), then verify each against evidence before forming any code-layer hypothesis.
Surgical changes — apply the diff test. Bug fix: every changed line must trace directly to the bug. Don't restyle or improve adjacent code. Enhancement task: implement improvements AND announce them explicitly.
Surface ambiguity before coding — don't pick silently. If request has multiple interpretations, present each with effort estimate and ask. Never assume all-records, file-based, or more complex path.

<!-- /SYNC:ai-mistake-prevention -->

<!-- SYNC:critical-thinking-mindset -->

Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.

<!-- /SYNC:critical-thinking-mindset -->

<!-- SYNC:critical-thinking-mindset:reminder -->

MUST ATTENTION apply critical thinking — every claim needs traced proof, confidence >80% to act. Anti-hallucination: never present guess as fact.

<!-- /SYNC:critical-thinking-mindset:reminder -->

<!-- SYNC:ai-mistake-prevention:reminder -->

MUST ATTENTION apply AI mistake prevention — holistic-first debugging, fix at responsible layer, surface ambiguity before coding, re-read files after compaction.

<!-- /SYNC:ai-mistake-prevention:reminder -->

Closing Reminders

  • MANDATORY IMPORTANT MUST ATTENTION break work into small todo tasks using task tracking BEFORE starting
  • MANDATORY IMPORTANT MUST ATTENTION search codebase for 3+ similar patterns before creating new code
  • MANDATORY IMPORTANT MUST ATTENTION cite file:line evidence for every claim (confidence >80% to act)
  • MANDATORY IMPORTANT MUST ATTENTION add a final review todo task to verify work quality

[TASK-PLANNING] Before acting, analyze task scope and systematically break it into small todo tasks and sub-tasks using task tracking.

<!-- CODEX:SYNC-PROMPT-PROTOCOLS:START -->

Hookless Prompt Protocol Mirror (Auto-Synced)

Source: .claude/hooks/lib/prompt-injections.cjs + .claude/.ck.json

[WORKFLOW-EXECUTION-PROTOCOL] [BLOCKING] Workflow Execution Protocol — MANDATORY IMPORTANT MUST CRITICAL. Do not skip for any reason.

Generic portability boundary: Reusable skills and protocol text stay project-neutral; project-specific conventions are discovered from docs/project-config.json and docs/project-reference/. Apply shared AI-SDD from shared/sdd-artifact-contract.md. Read docs/project-config.json and docs/project-reference/docs-index-reference.md, then open the project reference docs named there. Any supported AI tool may execute when this shared context and local docs are available.

1. DETECT: Match prompt against workflow catalog 2. ANALYZE: Find best-match workflow AND evaluate if a custom step combination would fit better 3. ASK (REQUIRED FORMAT): Use a direct user question with this structure unless the user explicitly invoked a workflow/skill and the local protocol treats explicit invocation as confirmation:

  • Question: "Which workflow do you want to activate?"
  • Option 1: "Activate [BestMatch Workflow] (Recommended)"
  • Option 2: "Activate custom workflow: [step1 → step2 → ...]" (include one-line rationale)

4. ACTIVATE (if confirmed): Call $workflow-start <workflowId> for standard; sequence custom steps manually 5. CREATE TASKS: task tracking for ALL workflow steps 6. EXECUTE: Follow each step in sequence [CRITICAL-THINKING-MINDSET] Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act. Anti-hallucination principle: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination. AI Attention principle (Primacy-Recency): Put the 3 most critical rules at both top and bottom of long prompts/protocols so instruction adherence survives long context windows. Goal-driven execution: Define success criteria first, loop until verified, and stop only when observable checks pass. Tests verify intent: Tests must protect business rules/invariants and fail when the protected intent breaks, not only mirror current behavior.

[LESSON-LEARNED-REMINDER] [BLOCKING] Task Planning & Continuous Improvement — MANDATORY. Do not skip.

Break work into small tasks (task tracking) before starting. Add final task: "Analyze AI mistakes & lessons learned".

Extract lessons — ROOT CAUSE ONLY, not symptom fixes:

1. Name the FAILURE MODE (reasoning/assumption failure), not symptom — "assumed API existed without reading source" not "used wrong enum value". 2. Generality test: does this failure mode apply to ≥3 contexts/codebases? If not, abstract one level up. 3. Write as a universal rule — strip project-specific names/paths/classes. Useful on any codebase. 4. Consolidate: multiple mistakes sharing one failure mode → ONE lesson. 5. Recurrence gate: "Would this recur in future session WITHOUT this reminder?" — No → skip $learn. 6. Auto-fix gate: "Could $code-review/$code-simplifier/$security/$lint catch this?" — Yes → improve review skill instead. 7. BOTH gates pass → ask user to run $learn. [TASK-PLANNING] [MANDATORY] BEFORE executing any workflow or skill step, create/update task tracking for all planned steps, then keep it synchronized as each step starts/completes.

<!-- CODEX:SYNC-PROMPT-PROTOCOLS:END -->

Related skills

AI & Agent Buildingllmautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.