Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
garrytan avatar

Media Ingest

  • 176 installs
  • 27.8k repo stars
  • Updated August 5, 2026
  • garrytan/gbrain

Ingest images, video, audio, and documents into gbrain so agents can index, summarize, and recall multimodal sources during research or content workflows.

About

media-ingest adds multimodal intake to garrytan/gbrain: images, audio, video, and documents are parsed, normalized, and stored so agents can search, cite, and reason over ingested assets instead of ephemeral chat context.

  • Multimodal file ingestion
  • Automated indexing pipeline
  • Agent-ready source normalization
  • Batch and single-asset intake
  • Feeds downstream query and reports

Media Ingest by the numbers

  • 176 all-time installs (skills.sh)
  • Ranked #3,085 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/garrytan/gbrain --skill media-ingest

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs176
repo stars27.8k
Last updatedAugust 5, 2026
Repositorygarrytan/gbrain

What it does

Ingest images, video, audio, and documents into gbrain so agents can index, summarize, and recall multimodal sources during research or content workflows.

Files

SKILL.mdMarkdownGitHub ↗

Media Ingest Skill

Ingest video, audio, PDF, book, screenshot, and GitHub repo content into the brain.

Filing rule: Read skills/_brain-filing-rules.md before creating any new page.

Contract

This skill guarantees:

  • Every ingested media item has a brain page with analysis (not just a transcript dump)
  • Transcripts (video/audio) saved in raw and human-readable formats
  • Entity extraction: every person and company mentioned gets back-linked
  • Raw source files preserved via gbrain files upload-raw
  • Filing by primary subject, not by media format
Convention: See skills/conventions/quality.md for Iron Law back-linking.

Every mention of a person or company with a brain page MUST create a back-link.

Phases

Phase 1: Identify format and fetch

FormatAction
YouTube/video URLFetch transcript (Whisper, transcription service, or captions)
Audio fileTranscribe with available STT service
PDFExtract text (OCR if needed)
Book PDFExtract text, identify chapters/sections
Screenshot/imageOCR via vision model, extract text and entities
GitHub repoClone, read README + key files, summarize architecture

Phase 2: Upload raw source

Save the original file for provenance: gbrain files upload-raw <file> --page <slug>

Phase 3: Create brain page

File by primary subject (not format). Use this template:

# {Title}

**Source:** {URL or file path}
**Format:** {video/audio/PDF/book/screenshot/repo}
**Created:** {date}

## Summary
{Key points, not a transcript dump}

## Key Segments / Highlights
{For video/audio: timestamped highlights. For books: chapter summaries.}

## People Mentioned
{List with links to brain pages}

## Companies Mentioned
{List with links to brain pages}

Phase 4: Entity extraction and propagation

For every person and company mentioned: 1. Check brain for existing page 2. Create/enrich if needed (delegate to enrich skill) 3. Add back-link from entity page to this media page 4. Add timeline entry on entity page

A media item is NOT fully ingested until entity propagation is complete.

Phase 5: Sync

gbrain sync to update the index.

Output Format

Brain page created with summary, highlights, and entity cross-links. Report to user: "Ingested {title}: {N} entities detected, {N} pages updated."

Anti-Patterns

  • Dumping raw transcripts without analysis
  • Skipping entity extraction ("I'll do that separately")
  • Filing raw ingest by format (all videos in media/videos/) instead of by subject. Note: format-prefixed paths under media/<format>/<slug> ARE sanctioned for synthesized one-of-one output like book-mirror's media/books/<slug>-personalized.md. The anti-pattern is for raw ingest, not for sui generis synthesis. See skills/_brain-filing-rules.md "Sanctioned exception: synthesis output is sui generis."
  • Not preserving raw source files
  • Creating stub pages without meaningful content

Related skills

AI & Agent Buildingautomationresearch

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.