Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
freestylefly avatar

Volcengine Video Understanding

  • 544 installs
  • 424 repo stars
  • Updated June 8, 2026
  • freestylefly/canghe-skills

volcengine-video-understanding is a Claude Code skill that lets coding agents understand, caption, and extract insights from video using Volcengine multimodal models for developers building video-aware automation feature

About

volcengine-video-understanding is a multimodal integration skill from freestylefly/canghe-skills, ranked 8 on Skills.sh with 445 installs. It equips coding agents to analyze video content through Volcengine's models—generating captions, summaries, and structured insights from footage. Developers invoke it when building features that ingest user uploads, surveillance clips, or marketing reels and need programmatic scene or content understanding without standing up a custom vision pipeline. The skill assumes Volcengine API access and fits agent workflows that combine video input with downstream code generation or data extraction.

  • Enables Claude, Cursor and other agents to process video files directly
  • Returns structured understanding including scene descriptions, objects, actions and timestamps
  • Supports both short clips and longer videos through intelligent chunking
  • Integrates as a reusable MCP-compatible tool in agent workflows
  • Handles video formats commonly used in product demos, tutorials and user-generated content

Volcengine Video Understanding by the numbers

  • 544 all-time installs (skills.sh)
  • Ranked #1,669 of 16,556 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Jul 31, 2026 (Skillselion catalog sync)
npx skills add https://github.com/freestylefly/canghe-skills --skill volcengine-video-understanding

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs544
repo stars424
Last updatedJune 8, 2026
Repositoryfreestylefly/canghe-skills

How do agents analyze video with Volcengine APIs?

Let their coding agent understand, caption, and extract insights from video content using Volcengine's multimodal models.

Who is it for?

Developers integrating Volcengine multimodal video APIs into agent workflows for captioning, summarization, or content analysis.

Skip if: Developers needing real-time video streaming infrastructure, custom model training, or non-Volcengine vision providers.

When should I use this skill?

A developer asks to understand, caption, summarize, or extract data from video files using Volcengine multimodal models.

What you get

Video captions, content summaries, and structured insight extracts from Volcengine multimodal responses.

  • Video captions
  • Content summaries
  • Structured insight JSON

By the numbers

  • 445 installs on Skills.sh
  • Rank 8 in freestylefly/canghe-skills catalog

Files

SKILL.mdMarkdownGitHub ↗

火山视频理解

使用字节跳动火山方舟视频理解 API(doubao-seed-2-0-pro-260215 等模型)对视频进行深度理解和分析。

推荐方式:Files API 上传 + Responses API 分析

  • 支持最大 512MB 视频文件
  • 自动视频预处理(FPS采样)
  • 文件可重复使用(存储7天)

功能

  • 视频上传:通过 Files API 上传本地视频(推荐,最大512MB)
  • 内容理解:分析视频场景、人物、动作、情感
  • 视频问答:基于视频内容回答用户问题
  • 视频描述:自动生成视频描述和摘要

前置要求

需要设置 ARK_API_KEY 环境变量。

配置方式(推荐)

1. 复制配置模板:

cp .canghe-skills/.env.example .canghe-skills/.env

2. 编辑 .canghe-skills/.env 文件,填写你的 API Key:

ARK_API_KEY=your-actual-api-key-here

或使用环境变量

export ARK_API_KEY="your-api-key"

加载优先级

1. 系统环境变量 (process.env) 2. 当前目录 .canghe-skills/.env 3. 用户主目录 ~/.canghe-skills/.env

使用方法

1. 基础视频分析(Files API 方式 - 推荐)

cd ~/.openclaw/workspace/skills/volcengine-video-understanding
python3 scripts/video_understand.py /path/to/video.mp4 "描述这个视频的内容"

2. 视频问答

python3 scripts/video_understand.py /path/to/video.mp4 "视频中出现了哪些人物?"

3. 情感分析

python3 scripts/video_understand.py /path/to/video.mp4 "分析视频中人物的情感变化"

4. 指定模型和帧率

python3 scripts/video_understand.py /path/to/video.mp4 "总结视频要点" \
  --model doubao-seed-2-0-pro-260215 \
  --fps 2

5. 保存结果到文件

python3 scripts/video_understand.py /path/to/video.mp4 "描述视频" --output result.json

参数说明

参数默认值说明
video_path必填视频文件路径
instruction必填分析指令/问题
--modeldoubao-seed-2-0-pro-260215模型 ID
--fps1视频采样帧率(预处理)
--output-结果输出文件路径

支持的模型

  • doubao-seed-2-0-pro-260215 (默认)
  • doubao-seed-2-0-lite-250728
  • doubao-seed-1-6-251015
  • 其他 Seed 系列视频理解模型

分析示例

示例 1:视频内容描述

python3 scripts/video_understand.py ~/Desktop/video.mp4 "详细描述这个视频的内容,包括场景、人物和动作"

示例 2:视频摘要

python3 scripts/video_understand.py ~/Desktop/video.mp4 "用3句话总结这个视频的要点"

示例 3:动作识别

python3 scripts/video_understand.py ~/Desktop/video.mp4 "视频中的人物在做什么动作?按时间顺序描述"

示例 4:场景分析

python3 scripts/video_understand.py ~/Desktop/video.mp4 "分析视频中的场景变化和环境特征"

技术细节

调用流程

1. 上传视频:通过 Files API 上传本地视频文件,指定 FPS 预处理配置 2. 等待处理:等待视频预处理完成(状态变为 processed) 3. 创建任务:调用 Responses API 进行视频理解 4. 获取结果:返回分析结果

API 格式

Files API 上传

curl https://ark.cn-beijing.volces.com/api/v3/files \
  -H "Authorization: Bearer $ARK_API_KEY" \
  -F 'purpose=user_data' \
  -F 'file=@video.mp4' \
  -F 'preprocess_configs[video][fps]=1'

Responses API 分析

{
  "model": "doubao-seed-2-0-pro-260215",
  "input": [
    {
      "role": "user",
      "content": [
        {
          "type": "input_video",
          "file_id": "file-xxxx"
        },
        {
          "type": "input_text",
          "text": "用户指令"
        }
      ]
    }
  ]
}

FPS 设置建议

FPS适用场景
0.3-0.5慢节奏视频、静态场景、节省token
1一般视频分析(默认)
2-3快速动作、细节分析

限制

  • 视频格式:MP4(推荐)、MOV、AVI
  • 文件大小:最大 512MB(Files API 方式)
  • 存储时间:上传的文件默认存储 7 天
  • 处理时间:根据视频长度和复杂度,通常 10-60 秒

Python API 使用

from scripts.video_understand import analyze_video

result = analyze_video(
    file_path="/path/to/video.mp4",
    instruction="描述视频内容",
    model="doubao-seed-2-0-pro-260215",
    fps=1
)

# 提取回答
text = ""
for item in result.get("output", []):
    if item.get("type") == "message":
        for content in item.get("content", []):
            if content.get("type") == "output_text":
                text = content.get("text", "")
                break

print(text)

错误处理

常见错误及解决方案:

错误原因解决方案
API Key 错误未设置或错误检查 ARK_API_KEY 环境变量
文件不存在路径错误检查文件路径
上传失败文件过大或格式不支持检查文件大小(<512MB)和格式
处理超时视频过长或复杂缩短视频或降低 FPS

参考文档

Related skills

FAQ

What does volcengine-video-understanding extract from video?

volcengine-video-understanding uses Volcengine multimodal models so coding agents can caption footage, summarize scenes, and return structured insights from video content.

How popular is volcengine-video-understanding on Skills.sh?

volcengine-video-understanding shows 445 installs and rank 8 on Skills.sh under freestylefly/canghe-skills among video AI integration skills.

AI & Agent Buildingagentsautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.