
Video Expert Analyzer
- 35 installs
- 229 repo stars
- Updated March 25, 2026
- albedo-tabai/video-expert-analyzer
video-expert-analyzer is a skill that scores and curates video scenes with a multimodal AI model using Walter Murch's editing rules and dynamic weighting.
About
video-expert-analyzer downloads a video, splits it into scenes, and scores each scene with a multimodal AI model using Walter Murch's six editing rules and a dynamic weighting system. It rates scenes on five dimensions and sorts them into MUST KEEP, USABLE, or DISCARD, copying the picks into a best_shots folder with a full analysis report. A developer or editor uses it to curate the strongest shots from source footage. It needs a vision-capable model (Gemini, Kimi, or Claude in Agent mode).
- Scores and curates video scenes with a multimodal AI model using Walter Murch's six editing rules and dynamic weighting
- Rates each scene on five dimensions and sorts into MUST KEEP, USABLE, or DISCARD, copying picks to scenes/best_shots/
- Supports Bilibili, YouTube, Douyin, and Xiaohongshu, with an Agent mode (host AI views frames) and an API mode
Video Expert Analyzer by the numbers
- 35 all-time installs (skills.sh)
- Ranked #946 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 2, 2026 (Skillselion catalog sync)
video-expert-analyzer capabilities & compatibility
Agent mode needs no API key (host AI views frames); API mode requires a VIDEO_ANALYZER_API_KEY for a vision model
- Capabilities
- video curation · scene scoring
- Works with
- openai
- Use cases
- video generation
- IDEs
- cursor ide · vscode
- Pricing
- Bring your own API key
What video-expert-analyzer says it does
基于 **Walter Murch 剪辑六法则** 和 **AI 自动评分系统** 的专业视频分析工具。
本 skill 的 AI 评分功能需要多模态(视觉理解)模型才能正常工作。
npx skills add https://github.com/albedo-tabai/video-expert-analyzer --skill video-expert-analyzerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 35 |
|---|---|
| repo stars | ★ 229 |
| Last updated | March 25, 2026 |
| Repository | albedo-tabai/video-expert-analyzer ↗ |
What it does
Score and curate video scenes with a multimodal AI model to extract the best shots from footage.
Who is it for?
Curating the strongest shots and highlights from source video footage
Skip if: Text-only models with no vision capability, which cannot score frames
When should I use this skill?
A user wants to analyze videos, score or rate scenes, extract best shots, or curate clips
What you get
Each scene is AI-scored on five dimensions and sorted into keep/usable/discard tiers with the best shots collected.
- scene_scores.json with per-scene ratings
- Best shots copied to scenes/best_shots/
- A complete analysis report
By the numbers
- 5 scoring dimensions (Aesthetic, Credibility, Impact, Memorability, Fun)
- 3 selection levels (MUST KEEP, USABLE, DISCARD)
- 4 scene types (Hook, Narrative, Aesthetic, Commercial)
Files
Video Expert Analyzer - 视频专家分析工具
基于 Walter Murch 剪辑六法则 和 AI 自动评分系统 的专业视频分析工具。
支持平台
| 平台 | 支持状态 | 说明 |
|---|---|---|
| Bilibili | ✅ 完全支持 | yt-dlp 下载 + B站API字幕 |
| YouTube | ✅ 完全支持 | yt-dlp 下载 |
| 抖音 (Douyin) | ✅ 完全支持 | 专用下载器(公开/分享链接无需浏览器 cookie) |
| 小红书 (Xiaohongshu) | ✅ 完全支持 | 专用下载器 |
| 其他平台 | ⚠️ 可能支持 | 取决于 yt-dlp 支持情况 |
核心特性
✅ 真实 AI 视觉评分 - 调用多模态大模型(Gemini 3.0/Kimi 2.5)真实分析画面内容 ✅ 双路径评分 - 支持「Agent 模式」(宿主 AI 直接看图)和「API 模式」(远程 API 调用) ✅ 中英双语术语 - 所有专业术语附中文释义 ✅ 可配置输出目录 - 首次使用设置,后续自动使用 ✅ 精选片段自动复制 - 自动复制到 scenes/best_shots/ ✅ 完整分析报告 - 包含理论依据、评分理由、整体评价(非模板) ✅ 动态权重系统 - 根据场景类型自动调整评分权重 ✅ 智能文件夹命名 - 以视频标题(自动裁剪)命名输出文件夹,更直观易懂
⚠️ 重要提示:本 skill 的 AI 评分功能需要多模态(视觉理解)模型才能正常工作。
请确保使用 Gemini 3.0、Kimi 2.5 等具备视觉能力的模型来调用此 skill。
模型兼容性
| 模型 | Agent 模式 | API 模式 | 说明 |
|---|---|---|---|
| Gemini 3.0 Flash | ✅ 推荐 | ✅ 推荐 | 速度快、视觉能力强 |
| Gemini 3.0 Pro | ✅ 推荐 | ✅ 支持 | 最强视觉理解 |
| Kimi 2.5 | ✅ 支持 | ✅ 支持 | 中文语境优秀 |
| Claude (Sonnet/Opus) | ✅ 支持 | ❌ 不支持 | 有视觉能力但无 OpenAI 兼容 API |
| 纯文本模型 | ❌ 不可用 | ❌ 不可用 | 无视觉能力,无法评分 |
Agent 模式 = 在 IDE(Cursor/VS Code/OpenClaw)中,AI 助手直接查看帧图片评分
API 模式 = 在终端 CLI 中,通过 OpenAI 兼容 API 远程调用视觉模型
术语对照表 (Terminology)
五维评分维度 (Five Dimensions)
| 英文术语 | 中文释义 | 详细说明 |
|---|---|---|
| Aesthetic Beauty | 美感 | 构图(三分法)、光影质感、色彩和谐度 |
| Credibility | 可信度 | 表演自然度、物理逻辑真实感、无出戏感 |
| Impact | 冲击力 | 视觉显著性(Visual Saliency)、动态张力、第一眼吸引力 |
| Memorability | 记忆度 | 独特视觉符号、冯·雷斯托夫效应(Von Restorff Effect)、金句 |
| Fun/Interest | 趣味度 | 参与感、娱乐价值、社交货币(Social Currency)潜力 |
场景类型 (Scene Types)
| 类型 | 中文释义 | 特点 |
|---|---|---|
| TYPE-A Hook | 钩子/开场型 | 高冲击力、吸引注意力 |
| TYPE-B Narrative | 叙事/情感型 | 人物对话、情感表达 |
| TYPE-C Aesthetic | 氛围/空镜型 | 风景、慢动作、极简构图 |
| TYPE-D Commercial | 商业/展示型 | 产品特写、广告展示 |
筛选等级 (Selection Levels)
| 等级 | 中文释义 | 标准 |
|---|---|---|
| MUST KEEP | 强烈推荐保留 | 加权总分 ≥ 8.5 或 单项 = 10 |
| USABLE | 可用素材 | 7.0 ≤ 加权总分 < 8.5 |
| DISCARD | 建议舍弃 | 加权总分 < 7.0 |
理论术语 (Theory Terms)
| 英文 | 中文 | 说明 |
|---|---|---|
| Walter Murch's Six Rules | 沃尔特·默奇六法则 | 情感>故事>节奏>视线追踪>2D平面>3D空间 |
| Von Restorff Effect | 冯·雷斯托夫效应 | 独特的项目更容易被记住 |
| Visual Saliency | 视觉显著性 | 吸引眼球的程度 |
| Social Currency | 社交货币 | 内容被分享的价值 |
| CTA | 行动号召 | Call to Action |
| SYNC | 节奏同步 | 画面与音频节拍的契合度 |
快速开始
第 1 步:数据处理 Pipeline
# 首次配置(只需一次)
python3 scripts/pipeline_enhanced.py --setup
# 分析视频
python3 scripts/pipeline_enhanced.py https://www.bilibili.com/video/BV1xxxxx
python3 scripts/pipeline_enhanced.py "https://www.douyin.com/video/xxxxx"第 2 步:AI 评分(二选一)
首次使用时,请选择评分模式:
🅰️ Agent 模式(推荐,IDE/OpenClaw/Cursor 用户)
如果你正在 IDE 或 AI 编程助手中使用此 skill,无需配置任何 API Key。 宿主 AI 助手(如 Gemini、Kimi)本身就具备视觉理解能力,可以直接「看图打分」。
流程: 当你(AI 助手)拥有视觉理解能力时,在 pipeline 完成后执行以下步骤:
1. 统计场景总数:先列出 <output_dir>/frames/ 目录,确定总共有多少个场景帧。这决定了后续分批策略。 2. 分批查看帧画面:每批 5-10 张帧图片为一个 batch(每批超过 10 张会降低单图分析质量,少于 5 张效率太低)。使用 view_file 工具批量查看 .jpg 文件。 3. 按以下维度为每个场景打分(1-10 整数):
- aesthetic_beauty(美感):构图、光影、色彩
- credibility(可信度):真实感、物理逻辑
- impact(冲击力):视觉显著性、第一眼吸引力
- memorability(记忆度):独特符号、冯·雷斯托夫效应
- fun_interest(趣味度):参与感、娱乐价值
4. 分类场景类型:TYPE-A Hook / TYPE-B Narrative / TYPE-C Aesthetic / TYPE-D Commercial 5. 计算加权分并筛选:加权 ≥ 8.5 → MUST KEEP,≥ 7.0 → USABLE,< 7.0 → DISCARD 6. 将评分结果更新到 `scene_scores.json` 中每个场景的字段 7. 将精选片段(MUST KEEP + USABLE)对应的 mp4 复制到 `scenes/best_shots/` 并按排名命名 8. 生成分析报告 <video_id>_complete_analysis.md
⚡ 大规模场景分批处理协议(>10 个场景时启用)
当场景数量超过 10 个时,必须使用分批分析模式。这不是建议,而是硬性要求——因为一次性处理过多场景时,AI 容易因上下文疲劳而跳过、抽样或给出模板化评分,导致分析质量崩塌。
🚫 反抽样声明(Anti-Sampling Policy)
每一个场景都是用户花了真金白银下载和切分的素材。抽样分析意味着用户会错过可能的最佳镜头——想象一下,被跳过的那个场景恰好是整个视频的高光时刻。因此:
- 禁止抽样、跳过、合并或"代表性选取"任何场景
- 禁止对未查看的场景给出评分(不看图就打分等于捏造数据)
- 禁止对多个场景使用相同的描述文字(每个镜头的画面内容不同,描述必然不同)
- 如果因为上下文限制确实无法继续,应当明确告知用户已完成到哪个场景,让用户决定如何继续,而不是悄悄跳过
分批规则:
- 每个 batch 包含 5-10 个场景(根据总数均匀划分)
- 按场景编号顺序处理:Batch 1 = Scene 001-010,Batch 2 = Scene 011-020,以此类推
每个 batch 执行流程: 1. 使用 view_file 查看本批次所有帧图片 2. 逐帧分析,每个场景必须包含:
- 📝 唯一视觉描述(1-2 句):描述该帧画面中实际看到的具体内容(人物/物体/动作/环境/色调)。这是证明你真正查看了图片的证据——如果描述不出画面内容,说明没有认真看图
- 🎯 五维评分(1-10 整数)+ 场景类型 + 加权分 + 筛选等级
3. 立即将本批次评分写入 scene_scores.json 4. 输出子报告 batch_N_report.md(N = 批次号),包含:
- 本批次场景编号范围
- 每个场景的评分摘要表格(含视觉描述列)
- 本批次 MUST KEEP / USABLE / DISCARD 数量统计
5. 向用户汇报进度:"✅ Batch N/M 完成(场景 XXX-YYY),MUST KEEP: X 个,USABLE: Y 个。继续处理下一批..."
全部 batch 完成后: 1. 🔍 覆盖率校验(必须执行):
- 列出
frames/目录中的所有.jpg文件数量 = 应分析总数 - 统计
scene_scores.json中已有评分的场景数量 = 实际完成数 - 输出校验结果:"覆盖率:实际完成数/应分析总数 = XX%"
- 覆盖率必须 = 100%。如果不足,列出遗漏的场景编号并补充分析
2. 汇总所有子报告,生成完整分析报告 <video_id>_complete_analysis.md 3. 复制精选片段到 scenes/best_shots/ 4. 确认最终报告输出完成后,删除所有 batch_N_report.md 子报告 5. 向用户输出最终摘要:总场景数、各等级分布、Top 5 精选镜头
🅱️ API 模式(独立 CLI 运行用户)
如果你直接在终端运行脚本,需要配置视觉大模型 API:
# 设置 API 密钥(必需)
export VIDEO_ANALYZER_API_KEY="your-api-key"
# 可选:自定义端点和模型
export VIDEO_ANALYZER_BASE_URL="https://generativelanguage.googleapis.com/v1beta/openai"
export VIDEO_ANALYZER_MODEL="gemini-2.0-flash"
# 运行 AI 分析
cd ~/Downloads/video-analysis/<视频文件夹>
python3 <skill_dir>/scripts/ai_analyzer.py scene_scores.json --mode api工作流程
┌─────────────────────────────────────────────────────────────┐
│ VIDEO EXPERT ANALYZER v2.0 │
└─────────────────────────────────────────────────────────────┘
第 1 阶段: 数据处理 (pipeline_enhanced.py)
1. 📥 下载视频 → video.mp4
2. 🎵 提取音频 → video.m4a
3. 🎞️ 场景检测 (detect-content) → scenes/*.mp4
4. 🎤 智能字幕提取 → video.srt
(B站API → 内嵌字幕 → RapidOCR → FunASR 四级降级)
5. 🖼️ 帧提取 → frames/*.jpg
6. 📊 生成评分模板 → scene_scores.json
第 2 阶段: AI 视觉评分 (选择一种模式)
┌────────────────────────┬────────────────────────┐
│ 🅰️ Agent 模式 │ 🅱️ API 模式 │
│ 宿主 AI 直接看图评分 │ 远程调用视觉大模型 │
│ 无需 API Key │ 需要 API Key │
│ IDE/OpenClaw/Cursor │ 独立 CLI 运行 │
└────────────────────────┴────────────────────────┘
↓
7. 🤖 真实视觉分析评分 → 基于画面内容的真实评分
8. 🧮 动态权重计算 → 根据类型计算加权得分
9. ⭐ 精选镜头筛选 → 复制到 scenes/best_shots/
10. 📄 生成完整报告 → *_complete_analysis.md分析方法论
Walter Murch 剪辑六法则
情感 Emotion > 故事 Story > 节奏 Rhythm > 视线追踪 Eye-trace > 2D平面 2D Plane > 3D空间 3D Space
一个情感真挚但画面略抖的镜头,优于一个画面完美但内容空洞的镜头。
五维评分体系 (Five-Dimension Scoring)
| 维度 (Dimension) | 基础权重 | 评估要点 |
|---|---|---|
| Aesthetic Beauty 美感 | 20% | 构图(三分法)、光影质感、色彩和谐度 |
| Credibility 可信度 | 20% | 表演自然度、物理逻辑、无出戏感 |
| Impact 冲击力 | 20% | 视觉显著性(Visual Saliency)、动态张力 |
| Memorability 记忆度 | 20% | 独特符号(Von Restorff Effect)、金句 |
| Fun/Interest 趣味度 | 20% | 参与感、娱乐价值、社交货币 |
动态权重系统(根据场景类型自动调整)
| 类型 (Type) | 权重分配 (Weighting) | 适用场景 (Application) |
|---|---|---|
| TYPE-A Hook | IMPACT 40% + MEMORABILITY 30% + SYNC 20% | 开场钩子、高能时刻 |
| TYPE-B Narrative | CREDIBILITY 40% + MEMORABILITY 30% + AESTHETICS 20% | 叙事段落、情感表达 |
| TYPE-C Aesthetic | AESTHETICS 50% + SYNC 30% + IMPACT 20% | 空镜头、氛围营造 |
| TYPE-D Commercial | CREDIBILITY 40% + MEMORABILITY 40% + AESTHETICS 20% | 产品展示、商业广告 |
筛选决策规则
| 等级 (Level) | 中文释义 | 标准 (Criteria) | 用途 (Usage) |
|---|---|---|---|
| 🌟 MUST KEEP | 强烈推荐保留 | 加权总分 > 8.5 或 单项 = 10 | 核心素材,极致长板 |
| 📁 USABLE | 可用素材 | 7.0 ≤ 加权总分 < 8.5 | 过渡素材,辅助叙事 |
| 🗑️ DISCARD | 建议舍弃 | 加权总分 < 7.0 或有瑕疵 | 建议舍弃 |
输出文件结构
{output_dir}/
└── {video_id}/
├── {video_id}.mp4 # 完整视频 (Full Video)
├── {video_id}.m4a # 音频文件 (Audio)
├── {video_id}.srt # 字幕文件 (Subtitles)
├── {video_id}_transcript.txt # 转录文本 (Transcript)
├── scene_scores.json # 完整评分数据(AI已填写)⭐
├── {video_id}_complete_analysis.md # ⭐ 完整分析报告(中英双语)
├── scenes/ # 场景片段 (Scene Clips)
│ ├── {video_id}-Scene-001.mp4
│ ├── {video_id}-Scene-002.mp4
│ ├── ...
│ └── best_shots/ # ⭐ 精选片段(已复制)
│ ├── 01_USABLE_xxx.mp4
│ ├── 02_MUST_KEEP_xxx.mp4
│ └── README.md # 双语说明文件
└── frames/ # 场景预览帧 (Preview Frames)
├── {video_id}-Scene-001.jpg
└── ...完整分析报告内容
生成的 *_complete_analysis.md 包含中英双语术语对照:
1. 统计概览 (Statistics Overview)
- 各等级场景数量统计(带中文释义)
- 各维度平均分(英文术语 + 中文释义)
2. 场景排名表 (Scene Rankings)
- 按加权分数排序
- 显示类型(英文 + 中文)
- 筛选建议(英文 + 中文)
- 核心优势
3. 术语对照表 (Terminology Reference)
完整的专业术语中英对照表,包含:
- 英文术语
- 中文释义
- 详细说明
4. 各场景详细评估(每个场景包含)
- 基础信息: 类型分类(双语)、加权得分、筛选建议(双语)
- 内容描述: AI 自动生成的场景描述
- 五维评分表格:
| 英文术语 | 中文释义 | 得分 | 权重贡献 |
- 入选/淘汰理由: AI 生成的理论依据
- 剪辑建议: AI 生成的使用建议
5. 精选片段推荐 (Best Shots Recommendations)
- 入选精选文件夹的片段列表(双语)
- 各类别最佳镜头:
- 最佳 Hook 候选 (Best Hook Candidate)
- 最佳视觉 (Best Visual)
- 最佳可信度 (Best Credibility)
- 最佳记忆度 (Best Memorability)
6. 整体影片评价 (Overall Assessment)
- 综合评分
- 评价结论(双语)
- 优势分析(各维度详细评价)
- 改进建议(专业术语 + 中文解释)
- 使用场景建议(社交媒体/产品展示/品牌宣传/广告投放)
使用示例
基础分析
# 配置输出目录(首次)
python3 scripts/pipeline_enhanced.py --setup
# 分析视频
python3 scripts/pipeline_enhanced.py https://www.bilibili.com/video/BV1xxxxx
# 进入输出目录运行 AI 分析
cd ~/Downloads/video-analysis/BV1xxxxx
python3 <skill_dir>/scripts/ai_analyzer.py scene_scores.json快速分析(自定义场景检测阈值)
python3 scripts/pipeline_enhanced.py URL --scene-threshold 20调整精选阈值
python3 scripts/ai_analyzer.py scene_scores.json 6.5 # 阈值 6.5命令参考
pipeline_enhanced.py
| 选项 | 说明 |
|---|---|
--setup | 配置输出目录 |
-o, --output | 指定输出目录 |
--scene-threshold | 场景检测阈值 (默认27) |
--best-threshold | 精选阈值 (默认7.5) |
ai_analyzer.py
| 参数 | 说明 |
|---|---|
scene_scores.json | 评分文件路径 |
--mode api | API 模式(需设置 VIDEO_ANALYZER_API_KEY) |
--mode agent | Agent 模式(生成模板供宿主 AI 填写) |
| 环境变量 | 说明 |
|---|---|
VIDEO_ANALYZER_API_KEY | 视觉大模型 API 密钥(API 模式必需) |
VIDEO_ANALYZER_BASE_URL | API 端点(默认 Gemini) |
VIDEO_ANALYZER_MODEL | 模型名称(默认 gemini-2.0-flash) |
依赖要求
# 系统依赖(安装 ffmpeg)
brew install ffmpeg # macOS
winget install ffmpeg # Windows
sudo apt install ffmpeg # Ubuntu/Debian
# 一键安装所有 Python 依赖
pip3 install -r requirements.txt
# 或手动安装核心依赖
pip3 install yt-dlp scenedetect[opencv] requests funasr modelscope torch torchaudio
# 可选依赖
pip3 install openai # API 模式评分
pip3 install rapidocr-onnxruntime # 烧录字幕 OCR 检测环境检测
# 一键检测所有依赖是否就绪
python3 scripts/check_environment.py平台特定说明
抖音视频下载
由于抖音的反爬机制,yt-dlp 无法直接下载抖音视频。本工具集成了专用的抖音下载器,可以:
- ✅ 自动识别抖音链接(支持
douyin.com和v.douyin.com短链接) - ✅ 自动提取视频信息(标题、作者)
- ✅ 下载无水印视频
- ✅ 自动提取音频用于转录
支持的抖音链接格式:
https://www.douyin.com/video/xxxxxhttps://www.douyin.com/jingxuan?modal_id=xxxxxhttps://v.douyin.com/xxxxx(短链接)
抖音下载实现原理: 1. 使用移动端 User-Agent 访问页面 2. 从页面 HTML 中提取 RENDER_DATA JSON 数据 3. 解析视频直链地址(自动替换 playwm 为 play 获取无水印版本) 4. 使用正确的 Referer 头下载视频
给 Agent 的硬性规则:
- 处理抖音链接时,不要先尝试网页登录、浏览器 cookie、WSL 读取宿主浏览器 cookie 这条路线
- 对公开可访问的抖音/分享链接,优先直接运行本地脚本:
python3 scripts/pipeline_enhanced.py "<抖音链接>"- 或
python3 scripts/download_douyin.py "<抖音链接>" ./video.mp4 - 如果 Agent 提示“抖音网页版需要登录”或“WSL 无法读取 cookie”,说明它走错路了,应切回本地下载脚本
- 优先使用用户从抖音 App 复制的分享短链(
https://v.douyin.com/...);长链也支持
Troubleshooting
终端命令卡顿
如果在 IDE 中执行 cp 或 Python 脚本时终端无响应,这通常是 IDE 终端代理的问题。解决方案:
- 方案 A: 打开 macOS 自带 Terminal.app 手动运行命令
- 方案 B: 使用
copy_by_index.py脚本通过 JSON 索引批量复制文件
FunASR 首次下载慢
FunASR 首次运行需下载约 2-3GB 的 Paraformer 模型。如果下载缓慢:
- 设置镜像:
export MODELSCOPE_CACHE=~/.cache/modelscope - 或使用 B站 API 字幕(无需本地模型,速度极快)
Agent 模式评分注意事项
- 必须使用具备视觉能力的多模态模型(参见「模型兼容性」表)
- 纯文本模型(如 GPT-3.5、Claude Haiku)无法执行 Agent 模式评分
- 建议每批查看 3-5 张帧图片,避免单次加载过多
抖音链接在 WSL / 远程环境报 cookie 错误
如果 Agent 在 WSL、远程容器或无桌面浏览器环境里说“无法读取抖音链接,因为网页版需要登录且 cookie 读不到”,按下面处理:
- 不要继续折腾网页登录
- 直接运行
scripts/pipeline_enhanced.py或scripts/download_douyin.py - 如果用户给的是长链,优先让用户从抖音 App 重新复制一次分享链接,拿到
v.douyin.com短链后再跑 - 只有私密、删除、地区受限或已失效内容,才可能真的无法直接下载
---
更新日志
v2.2.0 (2026-03-24)
- ⚡ 大规模场景分批处理协议:>10 个场景时强制分批(5-10个/batch),每批输出子报告 + 进度汇报,全部完成后汇总并清理
- 🚫 反抽样策略:禁止跳过/合并场景,强制输出唯一视觉描述作为查看证据
- 🔍 100% 覆盖率校验:完成后校验帧数 vs 已评分数,覆盖率必须 = 100%
- 📝 重写 description 触发词:增加 8 个中文关键词 + 小红书平台
- 🌐 补充 Windows/Linux 的 ffmpeg 安装说明
- 🗑️ 删除过时的 QUICKSTART.md、清理运行日志
- ⚠️ 模型兼容性发现:Kimi 2.5 在 >30 镜头时出现偷懒行为(疑似服务端工具调用次数限制);GPT 5.4 可完成 127 镜头连续分析(预热后 11 分 48 秒);Gemini/Opus 尚未测试
v2.1.0 (2026-02-27)
- ✅ 智能字幕提取:对齐 video-copy-analyzer,支持 B站API→内嵌→RapidOCR→FunASR 四级降级
- ✅ 新增小红书支持:集成
xiaohongshu_downloader.py - ✅ 新增 `requirements.txt`:一键安装所有依赖
- ✅ 新增模型兼容性矩阵:明确不同模型的适用模式
- ✅ 重写 `check_environment.py`:检测 v2.1 实际依赖
- ✅ 清理废弃文件:移除旧版 pipeline、Whisper 转录等6个废弃脚本
- ✅ 新增 Troubleshooting 章节
v2.0.0 (2026-02-27)
- ✅ 重写 AI 评分系统:移除假数据模拟,接入真实视觉大模型
- ✅ 双路径评分:Agent 模式(宿主 AI 直接看图)+ API 模式(远程调用)
- ✅ 修复场景检测:
detect-adaptive→detect-content,镜头切分更精准 - ✅ 将语音转录引擎从 Whisper 替换为 FunASR (Paraformer-zh)
- ✅ 中文场景转录速度大幅提升(10分钟音频约22秒完成)
- ✅ 移除
--whisper-model参数(FunASR 使用固定 paraformer-zh 模型) - ✅ 内置 VAD 自动分段 + 标点恢复
v1.4.0 (2026-02-06)
- ✅ 新增抖音视频下载支持
- ✅ 自动识别抖音链接并切换下载方式
- ✅ 支持抖音短链接和长链接
- ✅ 自动获取无水印视频
v1.3.0 (2026-02-05)
- ✅ 新增中英双语术语对照
- ✅ 所有专业术语附中文释义
- ✅ 报告中的术语表格双语显示
v1.2.0 (2026-02-05)
- ✅ 新增 AI 自动分析功能
- ✅ 动态权重评分系统
- ✅ 自动生成完整分析报告
- ✅ 自动复制精选镜头
v1.1.0 (2026-02-05)
- ✅ 新增可配置输出目录功能
- ✅ 精选片段自动保存
- ✅ 生成详细分析报告模板
v1.0.0 (2026-02-05)
- 初始版本
- 基础视频分析功能
---
基于 Walter Murch 剪辑六法则 (Walter Murch's Six Rules of Editing) AI 自动评分系统 (AI Automatic Scoring System) 动态权重算法 (Dynamic Weighting Algorithm) 中英双语术语对照 (Bilingual Terminology Reference)
# Video Expert Analyzer v2.1 配置示例
# 将此文件复制到输出目录并按需修改
# === AI 评分模式 ===
# API 模式需要配置以下变量(Agent 模式无需配置)
VIDEO_ANALYZER_API_KEY=your-api-key-here
VIDEO_ANALYZER_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
VIDEO_ANALYZER_MODEL=gemini-2.0-flash
# === 场景检测 ===
SCENE_THRESHOLD=27.0
# === 精选阈值 ===
BEST_SHOT_THRESHOLD=7.5
# === 输出选项 ===
GENERATE_FRAMES=true
GENERATE_TRANSCRIPT=true
GENERATE_REPORT=true
# Python
__pycache__/
*.py[cod]
*$py.class
*.so
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
*.egg-info/
.installed.cfg
*.egg
# Virtual environments
.env
.venv
env/
venv/
ENV/
# IDE
.idea/
.vscode/
*.swp
*.swo
*~
# OS
.DS_Store
.DS_Store?
._*
.Spotlight-V100
.Trashes
ehthumbs.db
Thumbs.db
# Output
output/
*.mp4
*.m4a
*.srt
*.wav
# Runtime logs
*.log
log.txt
# Temporary batch reports
batch_*_report.md
Changelog
All notable changes to the Video Expert Analyzer skill will be documented in this file.
[2.2.0] - 2026-03-24
Added
- ⚡ 大规模场景分批处理协议:>10 个场景时强制分批(5-10个/batch),每批输出子报告并汇报进度,全部完成后汇总为完整报告并清理子报告
- 🚫 反抽样策略(Anti-Sampling Policy):禁止抽样/跳过/合并场景,每个场景必须附带唯一视觉描述作为查看证据
- 🔍 100% 覆盖率校验:完成全部分析后校验 frames/ 中帧数与 scene_scores.json 已评分数,覆盖率必须 = 100%
- 🌐 补充 Windows (
winget) 和 Linux (apt) 的 ffmpeg 安装说明
Changed
- 📝 重写 description 触发词:增加中文关键词(视频分析/镜头筛选/场景评分/视频拆解/精选片段/镜头打分/素材挑选)和场景描述,提升 skill 触发准确率
- 补充小红书(Xiaohongshu)到 description 支持平台列表
- 硬编码路径
~/.openclaw/...替换为通用<skill_dir>
Removed
- 🗑️ 删除严重过时的
QUICKSTART.md(引用 v1.x 废弃的 pipeline.py / scoring_helper.py / Whisper 参数) - 清理运行日志
error.log、log.txt
Fixed
.gitignore补充*.log、log.txt、batch_*_report.md规则
Model Compatibility Notes
- ⚠️ Kimi 2.5:>30 个镜头时会出现偷懒行为(抽样分析),推测为服务端限制了模型最大连续工作时长或工具调用次数
- ✅ GPT 5.4:经测试可完成高达 127 个镜头的连续分析任务,模型预热后仅需 11 分 48 秒
- 🔲 Gemini / Opus:尚未测试大规模场景分析
---
[2.1.0] - 2026-02-27
Added
- ✅ 智能字幕提取:对齐 video-copy-analyzer,支持 B站API → 内嵌字幕 → RapidOCR → FunASR 四级降级
- ✅ 新增小红书支持:集成
xiaohongshu_downloader.py - ✅ 新增
requirements.txt:一键安装所有依赖 - ✅ 新增模型兼容性矩阵
- ✅ 重写
check_environment.py - ✅ 新增 Troubleshooting 章节
Removed
- 移除旧版 pipeline、Whisper 转录等 6 个废弃脚本
---
[2.0.0] - 2026-02-27
Changed
- ✅ 重写 AI 评分系统:移除假数据模拟,接入真实视觉大模型
- ✅ 双路径评分:Agent 模式(宿主 AI 直接看图)+ API 模式(远程调用)
- ✅ 修复场景检测:
detect-adaptive→detect-content - ✅ 语音转录引擎 Whisper → FunASR (Paraformer-zh)
- ✅ 内置 VAD 自动分段 + 标点恢复
---
[1.4.0] - 2026-02-06
Added
- 新增抖音视频下载支持(自动识别链接、无水印下载)
[1.3.0] - 2026-02-05
Added
- 新增中英双语术语对照
[1.2.0] - 2026-02-05
Added
- AI 自动分析功能、动态权重评分系统、自动生成完整分析报告
[1.1.0] - 2026-02-05
Added
- 可配置输出目录、精选片段自动保存
[1.0.0] - 2026-02-05
Added
- 初始版本:基于 Walter Murch 六法则的视频分析工具
- yt-dlp 下载 + PySceneDetect 场景检测 + Whisper 转录
- 五维评分体系、评分辅助工具、报告模板
<p align="center"> <img src="https://img.shields.io/badge/version-2.1.0-blue" alt="Version"> <img src="https://img.shields.io/badge/license-MIT-green" alt="License"> <img src="https://img.shields.io/badge/python-3.9+-yellow" alt="Python"> <img src="https://img.shields.io/badge/AI-Gemini%203.0%20%7C%20Kimi%202.5%20%7C%20Claude-purple" alt="AI Models"> </p>
<p align="center"> <b>🌐 Language / 语言</b><br> <a href="#english">English</a> | <a href="#chinese">中文</a> </p>
---
<a name="english"></a>
🎬 Video Expert Analyzer
AI-powered professional video analysis tool based on Walter Murch's Six Rules of Editing, with real multimodal AI vision scoring (Gemini 3.0 / Kimi 2.5 / Claude).
✨ Features
- 🤖 Real AI Vision Scoring — Multimodal models (Gemini/Kimi/GPT-4o) analyze actual frame content
- 🔀 Dual Scoring Paths — Agent mode (IDE AI reads frames) + API mode (remote API calls)
- 🎯 Dynamic Weighting — Weights auto-adjust based on scene type (Hook/Narrative/Aesthetic/Commercial)
- 🎬 Scene Detection — PySceneDetect
detect-contentfor accurate scene splitting - 🎤 Smart Subtitle Extraction — 4-tier fallback: Bilibili API → Embedded → RapidOCR → FunASR
- ⭐ Best Shots — Auto-copy top-rated clips to
best_shots/ - 📊 5D Scoring — Aesthetic, Credibility, Impact, Memorability, Fun/Interest
- 🌐 Bilingual — All terminology with Chinese translations
📱 Supported Platforms
| Platform | Status | Notes |
|---|---|---|
| Bilibili | ✅ Full Support | yt-dlp download + Bilibili API subtitles |
| YouTube | ✅ Full Support | yt-dlp download |
| Douyin (抖音) | ✅ Full Support | Dedicated downloader (public/share links do not need browser cookies) |
| Xiaohongshu (小红书) | ✅ Full Support | Dedicated downloader |
| Others | ⚠️ May Work | Depends on yt-dlp support |
🤖 Model Compatibility
| Model | Agent Mode | API Mode | Notes |
|---|---|---|---|
| Gemini 3.0 Flash | ✅ Recommended | ✅ Recommended | Fast, strong vision |
| Gemini 3.0 Pro | ✅ Recommended | ✅ Supported | Best visual understanding |
| Kimi 2.5 | ✅ Supported | ✅ Supported | Excellent for Chinese |
| Claude (Sonnet/Opus) | ✅ Supported | ❌ No | Has vision but no OpenAI-compatible API |
| Text-only models | ❌ No | ❌ No | Cannot score without vision |
Agent Mode = AI assistant in IDE views frame images directly
API Mode = CLI calls vision model via OpenAI-compatible API
🚀 Quick Start
Prerequisites
# System dependencies
brew install ffmpeg # macOS
# Install all Python dependencies
pip3 install -r requirements.txt
# Check environment
python3 scripts/check_environment.pyOne-Command Analysis
# Setup (first time only)
python3 scripts/pipeline_enhanced.py --setup
# Analyze any video
python3 scripts/pipeline_enhanced.py https://www.bilibili.com/video/BV1xxxxx
python3 scripts/pipeline_enhanced.py "https://www.douyin.com/video/xxxxx"
# AI scoring (choose one)
# Option A: Agent mode (in IDE, AI assistant scores visually)
# Option B: API mode
export VIDEO_ANALYZER_API_KEY="your-key"
python3 scripts/ai_analyzer.py scene_scores.json --mode api📊 Scoring System
Five Dimensions
| Dimension | Weight | Description |
|---|---|---|
| Aesthetic Beauty 美感 | 20% | Composition, lighting, color harmony |
| Credibility 可信度 | 20% | Authenticity, natural performance |
| Impact 冲击力 | 20% | Visual saliency, attention-grabbing |
| Memorability 记忆度 | 20% | Uniqueness, Von Restorff Effect |
| Fun/Interest 趣味度 | 20% | Engagement, entertainment, social currency |
Scene Types & Dynamic Weights
| Type | Primary Weights | Use Cases |
|---|---|---|
| TYPE-A Hook | Impact 40% + Memorability 30% | Opening hooks, high-energy moments |
| TYPE-B Narrative | Credibility 40% + Memorability 30% | Story segments, emotional scenes |
| TYPE-C Aesthetic | Aesthetics 50% + Sync 30% | B-roll, atmosphere shots |
| TYPE-D Commercial | Credibility 40% + Memorability 40% | Product showcases, ads |
Selection Levels
| Level | Criteria | Usage |
|---|---|---|
| 🌟 MUST KEEP | Score ≥ 8.5 or any dimension = 10 | Core material |
| 📁 USABLE | 7.0 ≤ Score < 8.5 | Supporting shots |
| 🗑️ DISCARD | Score < 7.0 | Not recommended |
📁 Output Structure
output-directory/
├── {video_id}.mp4 # Full video
├── {video_id}.m4a # Audio
├── {video_id}.srt # Subtitles (smart extraction)
├── scene_scores.json # ⭐ AI scoring data
├── *_complete_analysis.md # ⭐ Full analysis report
├── scenes/ # Scene clips
│ └── best_shots/ # ⭐ Top-rated clips (auto-copied)
└── frames/ # Preview frames🔧 Configuration
Pipeline Options
| Option | Description |
|---|---|
--setup | Configure output directory |
--scene-threshold | Scene detection sensitivity (default: 27) |
--best-threshold | Best shots threshold (default: 7.5) |
-o, --output | Output directory |
API Environment Variables
| Variable | Description |
|---|---|
VIDEO_ANALYZER_API_KEY | Vision model API key (required for API mode) |
VIDEO_ANALYZER_BASE_URL | API endpoint (default: Gemini) |
VIDEO_ANALYZER_MODEL | Model name (default: gemini-2.0-flash) |
📚 Theory Background
Based on Walter Murch's Six Rules:
Emotion > Story > Rhythm > Eye-trace > 2D Plane > 3D Space
A shot with genuine emotion but slight shake is better than a perfect but empty frame.
🙏 Credits
Built with:
- yt-dlp — Video download
- FunASR — Chinese speech recognition
- PySceneDetect — Scene detection
- FFmpeg — Media processing
- RapidOCR — Burned subtitle OCR
📖 References
Core Theory
1. Murch, W. (2001). In the Blink of an Eye (2nd ed.). Silman-James Press. 2. Murch, W. (1995). The Conversations. Knopf.
Psychology & Cognitive Science
3. Von Restorff, H. (1933). Psychologische Forschung, 18(1), 299-342. 4. Itti, L., & Koch, C. (2001). Nature Reviews Neuroscience, 2(3), 194-203. 5. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
Social Media & Virality
6. Berger, J. (2013). Contagious. Simon & Schuster. 7. Berger, J., & Milkman, K. L. (2012). Journal of Marketing Research, 49(2), 192-205.
Video & Film Analysis
8. Bordwell, D., & Thompson, K. (2012). Film Art (10th ed.). McGraw-Hill. 9. Katz, S. D. (1991). Film Directing Shot by Shot. Michael Wiese Productions. 10. Brown, B. (2016). Cinematography: Theory and Practice (3rd ed.). Routledge.
---
<a name="chinese"></a>
🎬 视频专家分析器
基于 Walter Murch 剪辑六法则 和 真实 AI 视觉评分 的专业视频分析工具
✨ 核心特性
- 🤖 真实 AI 视觉评分 — 多模态大模型(Gemini/Kimi/GPT-4o)分析真实画面内容
- 🔀 双路径评分 — Agent 模式(IDE 中 AI 直接看图)+ API 模式(远程 API 调用)
- 🎯 动态权重系统 — 根据场景类型自动调整权重(Hook/叙事/氛围/商业)
- 🎬 场景检测 — PySceneDetect
detect-content精准场景分割 - 🎤 智能字幕提取 — 四级降级:B站API → 内嵌字幕 → RapidOCR → FunASR
- ⭐ 精选片段 — 自动复制高分片段到
best_shots/ - 📊 五维评分 — 美感、可信度、冲击力、记忆度、趣味度
- 🌐 中英双语 — 所有专业术语附中文释义
📱 支持平台
| 平台 | 支持状态 | 说明 |
|---|---|---|
| Bilibili | ✅ 完全支持 | yt-dlp 下载 + B站API字幕 |
| YouTube | ✅ 完全支持 | yt-dlp 下载 |
| 抖音 (Douyin) | ✅ 完全支持 | 专用下载器(公开/分享链接无需浏览器 cookie) |
| 小红书 (Xiaohongshu) | ✅ 完全支持 | 专用下载器 |
| 其他平台 | ⚠️ 可能支持 | 取决于 yt-dlp |
🤖 模型兼容性
| 模型 | Agent 模式 | API 模式 | 说明 |
|---|---|---|---|
| Gemini 3.0 Flash | ✅ 推荐 | ✅ 推荐 | 速度快、视觉能力强 |
| Gemini 3.0 Pro | ✅ 推荐 | ✅ 支持 | 最强视觉理解 |
| Kimi 2.5 | ✅ 支持 | ✅ 支持 | 中文语境优秀 |
| Claude (Sonnet/Opus) | ✅ 支持 | ❌ 不支持 | 有视觉能力但无 OpenAI 兼容 API |
| 纯文本模型 | ❌ 不可用 | ❌ 不可用 | 无视觉能力 |
🚀 快速开始
环境准备
# 系统依赖
brew install ffmpeg
# 一键安装所有依赖
pip3 install -r requirements.txt
# 检查环境
python3 scripts/check_environment.py一键分析
# 首次配置
python3 scripts/pipeline_enhanced.py --setup
# 分析视频
python3 scripts/pipeline_enhanced.py https://www.bilibili.com/video/BV1xxxxx
python3 scripts/pipeline_enhanced.py "https://www.douyin.com/video/xxxxx"
# AI 评分(二选一)
# 方式 A:Agent 模式(IDE 中 AI 助手直接看图评分)
# 方式 B:API 模式
export VIDEO_ANALYZER_API_KEY="your-key"
python3 scripts/ai_analyzer.py scene_scores.json --mode api抖音链接说明
- 公开可访问的抖音链接,优先直接交给本地脚本处理,不要先尝试网页登录或浏览器 cookie
- 推荐命令:
python3 scripts/pipeline_enhanced.py "<抖音链接>"
# 或只下载视频
python3 scripts/download_douyin.py "<抖音链接>" ./video.mp4- 如果在 WSL、远程容器或无桌面浏览器环境里看到“cookie 无法读取”的报错,说明走错路线了,应切回上面的本地脚本
- 优先使用从抖音 App 复制的
https://v.douyin.com/...分享短链;长链同样支持
📊 评分体系
五维评分维度
| 维度 | 权重 | 评估要点 |
|---|---|---|
| 美感 (Aesthetic) | 20% | 构图(三分法)、光影质感、色彩和谐度 |
| 可信度 (Credibility) | 20% | 表演自然度、物理逻辑、无出戏感 |
| 冲击力 (Impact) | 20% | 视觉显著性、动态张力、第一眼吸引力 |
| 记忆度 (Memorability) | 20% | 独特视觉符号、冯·雷斯托夫效应 |
| 趣味度 (Fun) | 20% | 参与感、娱乐价值、社交货币潜力 |
筛选等级
| 等级 | 标准 | 用途 |
|---|---|---|
| 🌟 MUST KEEP | 加权总分 ≥ 8.5 或 单项 = 10 | 核心素材 |
| 📁 USABLE | 7.0 ≤ 加权总分 < 8.5 | 辅助素材 |
| 🗑️ DISCARD | 加权总分 < 7.0 | 建议舍弃 |
📁 输出结构
输出目录/
├── {video_id}.mp4 # 完整视频
├── {video_id}.m4a # 音频
├── {video_id}.srt # 字幕(智能提取)
├── scene_scores.json # ⭐ AI 评分数据
├── *_complete_analysis.md # ⭐ 完整分析报告
├── scenes/ # 场景片段
│ └── best_shots/ # ⭐ 精选片段(自动复制)
└── frames/ # 预览帧🔧 配置选项
| 选项 | 说明 |
|---|---|
--setup | 配置输出目录 |
--scene-threshold | 场景检测阈值 (默认: 27) |
--best-threshold | 精选阈值 (默认: 7.5) |
📚 理论背景
基于 Walter Murch 剪辑六法则:
情感 > 故事 > 节奏 > 视线追踪 > 2D平面 > 3D空间
一个情感真挚但画面略抖的镜头,优于一个画面完美但内容空洞的镜头。
🙏 致谢
- yt-dlp — 视频下载
- FunASR — 中文语音识别
- PySceneDetect — 场景检测
- FFmpeg — 媒体处理
- RapidOCR — 烧录字幕识别
---
📜 License
MIT License - 自由使用和修改
v2.1.0 Release Notes
🚀 What's New
🎤 Smart Subtitle Extraction (智能字幕提取)
Aligned with video-copy-analyzer skill — now supports 4-tier fallback: 1. Bilibili API — Fastest, grabs platform subtitles directly 2. Embedded Subtitles — Extracts subtitle streams via FFmpeg 3. Burned Subtitle OCR — Detects and extracts burned-in subtitles via RapidOCR 4. FunASR Speech-to-Text — Local Chinese ASR with VAD + punctuation recovery
🤖 Model Compatibility Matrix (模型兼容性矩阵)
Clear guidance on which AI models work with each scoring mode:
| Model | Agent Mode | API Mode |
|---|---|---|
| Gemini 3.0 Flash | ✅ Recommended | ✅ Recommended |
| Gemini 3.0 Pro | ✅ Recommended | ✅ Supported |
| Kimi 2.5 | ✅ Supported | ✅ Supported |
| Claude (Sonnet/Opus) | ✅ Supported | ❌ |
📱 Xiaohongshu Support (小红书支持)
Added xiaohongshu_downloader.py for downloading videos from Xiaohongshu (Little Red Book).
📦 One-Click Dependencies (一键安装依赖)
New requirements.txt — install everything with:
pip3 install -r requirements.txt🔍 Environment Check (环境检测)
Completely rewritten check_environment.py — now detects all v2.1 dependencies including FunASR, scenedetect, Apple MPS GPU support, etc.
🛠️ Troubleshooting Section
New troubleshooting guide in SKILL.md covering:
- IDE terminal freezing workarounds
- FunASR model download issues
- Agent mode scoring best practices
🧹 Cleanup
Removed 6 deprecated files:
scripts/pipeline.py(replaced bypipeline_enhanced.py)scripts/scoring_helper.py(replaced byscoring_helper_enhanced.py)scripts/transcribe_audio.py(Whisper, replaced by FunASR)scripts/douyin_downloader.py(replaced bydownload_douyin.py)xhs_debug.html(debug artifact)analyze.py(unused entry point)
📝 Changed Files
| File | Change |
|---|---|
SKILL.md | +Model compatibility +Xiaohongshu +Troubleshooting +v2.1 changelog |
README.md | Full rewrite from v1.3 to v2.1 |
requirements.txt | NEW — all dependencies |
.env.example | Rewritten (removed Whisper config) |
scripts/check_environment.py | Rewritten for v2.1 deps |
scripts/pipeline_enhanced.py | Transcription → smart subtitle extraction |
scripts/extract_subtitle_funasr.py | NEW — 4-tier subtitle extraction |
scripts/fetch_bilibili_subtitle.py | NEW — Bilibili API subtitles |
scripts/download_douyin.py | NEW — cleaner Douyin downloader |
⬆️ Upgrade from v2.0
pip3 install -r requirements.txt
python3 scripts/check_environment.pyFull Changelog: v2.0.0...v2.1.0
yt-dlp
scenedetect[opencv]
requests
funasr
modelscope
torch
torchaudio
openai
rapidocr-onnxruntime
#!/usr/bin/env python3
"""
AI Scene Analyzer v2.0
支持两种评分模式:
- API 模式:通过 OpenAI 兼容 API 调用 Gemini/Kimi 等视觉大模型
- Agent 模式:由宿主 AI 助手(IDE/OpenClaw 中的多模态模型)直接看图评分
环境变量(API 模式):
VIDEO_ANALYZER_API_KEY - API 密钥
VIDEO_ANALYZER_BASE_URL - API 端点 (默认 https://generativelanguage.googleapis.com/v1beta/openai)
VIDEO_ANALYZER_MODEL - 模型名称 (默认 gemini-2.0-flash)
"""
import json
import base64
import os
import re
import shutil
import sys
import time
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
# ============================================================
# 术语对照表
# ============================================================
TERMINOLOGY = {
"TYPE-A Hook": "TYPE-A Hook (钩子/开场型)",
"TYPE-B Narrative": "TYPE-B Narrative (叙事/情感型)",
"TYPE-C Aesthetic": "TYPE-C Aesthetic (氛围/空镜型)",
"TYPE-D Commercial": "TYPE-D Commercial (商业/展示型)",
"aesthetic_beauty": "美感 Aesthetic Beauty (构图/光影/色彩)",
"credibility": "可信度 Credibility (真实感/表演自然度)",
"impact": "冲击力 Impact (视觉显著性/动态张力)",
"memorability": "记忆度 Memorability (独特符号/金句)",
"fun_interest": "趣味度 Fun/Interest (参与感/娱乐价值)",
"MUST KEEP": "MUST KEEP (强烈推荐保留)",
"USABLE": "USABLE (可用素材)",
"DISCARD": "DISCARD (建议舍弃)",
}
def get_term_chinese(term: str) -> str:
return TERMINOLOGY.get(term, term)
# ============================================================
# System Prompt for Vision LLM
# ============================================================
SCORING_SYSTEM_PROMPT = """你是一位专业的视频剪辑/镜头分析专家,精通 Walter Murch 剪辑六法则。
你需要分析一张视频截帧,按以下维度打分(1-10 整数):
1. **aesthetic_beauty** (美感): 构图(如三分法/对称)、光影质感、色彩和谐度
2. **credibility** (可信度): 画面真实感、物理逻辑、AI生成痕迹程度(痕迹越少分越高)
3. **impact** (冲击力): 第一眼视觉显著性、动态张力、能否瞬间吸引观众
4. **memorability** (记忆度): 独特视觉符号、冯·雷斯托夫效应、过目不忘程度
5. **fun_interest** (趣味度): 参与感、娱乐价值、社交货币潜力
同时判断场景类型:
- TYPE-A Hook: 高冲击力开场/高能片段
- TYPE-B Narrative: 叙事/情感表达
- TYPE-C Aesthetic: 空镜/氛围/纯美学
- TYPE-D Commercial: 产品展示/商业广告
你必须严格按以下 JSON 格式返回结果,不要附加任何其他文字:
```json
{
"type_classification": "TYPE-X ...",
"description": "一句话中文描述画面内容",
"visual_summary": "视觉元素概要",
"scores": {
"aesthetic_beauty": 8,
"credibility": 7,
"impact": 9,
"memorability": 8,
"fun_interest": 7
},
"selection_reasoning": "入选/淘汰理由(中文)",
"edit_suggestion": "剪辑建议(中文)"
}
```"""
SCORING_USER_PROMPT_TEMPLATE = """请分析以下视频的第 {scene_num} 个场景截帧。
视频标题:{video_title}
视频总场景数:{total_scenes}
{transcript_info}
请严格按 JSON 格式返回分析结果。"""
# ============================================================
# API 模式:调用远程视觉大模型
# ============================================================
def call_vision_api(
frame_path: Path,
scene_num: int,
video_title: str = "",
total_scenes: int = 0,
transcript_text: str = "",
api_key: str = "",
base_url: str = "",
model: str = "",
) -> Optional[Dict]:
"""通过 OpenAI 兼容 API 调用视觉大模型分析帧画面"""
try:
from openai import OpenAI
except ImportError:
print(" ⚠️ 需要安装 openai 库: pip install openai")
return None
if not api_key:
print(" ⚠️ 未设置 VIDEO_ANALYZER_API_KEY 环境变量")
return None
# 读取图片并转 base64
with open(frame_path, "rb") as f:
image_data = base64.b64encode(f.read()).decode("utf-8")
mime_type = "image/jpeg" if frame_path.suffix.lower() in [".jpg", ".jpeg"] else "image/png"
transcript_info = f"对应转录文本片段:{transcript_text}" if transcript_text else "(该场景无转录文本)"
user_prompt = SCORING_USER_PROMPT_TEMPLATE.format(
scene_num=scene_num,
video_title=video_title,
total_scenes=total_scenes,
transcript_info=transcript_info,
)
client = OpenAI(api_key=api_key, base_url=base_url)
try:
response = client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": SCORING_SYSTEM_PROMPT},
{
"role": "user",
"content": [
{"type": "text", "text": user_prompt},
{
"type": "image_url",
"image_url": {
"url": f"data:{mime_type};base64,{image_data}",
},
},
],
},
],
temperature=0.3,
max_tokens=1024,
)
content = response.choices[0].message.content.strip()
# 提取 JSON(可能包裹在 ```json ... ``` 中)
json_match = re.search(r"\{[\s\S]*\}", content)
if json_match:
return json.loads(json_match.group())
else:
print(f" ⚠️ Scene {scene_num:03d}: 无法解析模型返回的 JSON")
return None
except Exception as e:
print(f" ⚠️ Scene {scene_num:03d}: API 调用失败 - {e}")
return None
# ============================================================
# 加权评分计算
# ============================================================
def compute_weighted_score(analysis: Dict) -> Dict:
"""根据场景类型计算加权分数并确定筛选等级"""
scores = analysis.get("scores", {})
type_class = analysis.get("type_classification", "")
# 根据类型动态调整权重
if "TYPE-A" in type_class:
weighted = (
scores.get("impact", 5) * 0.40
+ scores.get("memorability", 5) * 0.30
+ scores.get("aesthetic_beauty", 5) * 0.20
+ scores.get("fun_interest", 5) * 0.10
)
elif "TYPE-B" in type_class:
weighted = (
scores.get("credibility", 5) * 0.40
+ scores.get("memorability", 5) * 0.30
+ scores.get("aesthetic_beauty", 5) * 0.20
+ scores.get("fun_interest", 5) * 0.10
)
elif "TYPE-C" in type_class:
weighted = (
scores.get("aesthetic_beauty", 5) * 0.50
+ scores.get("impact", 5) * 0.20
+ scores.get("memorability", 5) * 0.20
+ scores.get("credibility", 5) * 0.10
)
else: # TYPE-D Commercial
weighted = (
scores.get("credibility", 5) * 0.40
+ scores.get("memorability", 5) * 0.40
+ scores.get("aesthetic_beauty", 5) * 0.20
)
analysis["weighted_score"] = round(weighted, 2)
# 确定筛选等级
if weighted >= 8.5 or any(v == 10 for v in scores.values()):
analysis["selection"] = "[MUST KEEP]"
elif weighted >= 7.0:
analysis["selection"] = "[USABLE]"
else:
analysis["selection"] = "[DISCARD]"
return analysis
# ============================================================
# 主流程:自动评分
# ============================================================
def auto_score_scenes(scores_path: Path, video_analysis_dir: Path, mode: str = "api") -> Dict:
"""自动为所有场景评分。mode='api' 使用远程API,mode='agent' 仅生成模板供宿主AI填写"""
with open(scores_path, "r", encoding="utf-8") as f:
data = json.load(f)
scenes = data.get("scenes", [])
frames_dir = video_analysis_dir / "frames"
video_title = data.get("title", data.get("video_id", ""))
total_scenes = len(scenes)
# 读取转录文本
transcript_text = ""
for ext in ["_transcript.txt", ".srt"]:
transcript_file = video_analysis_dir / f"{data.get('video_id', '')}{ext}"
if transcript_file.exists():
transcript_text = transcript_file.read_text(encoding="utf-8")
break
if mode == "agent":
print(f"\n📋 Agent 模式:已准备 {total_scenes} 个场景的评分模板")
print(f" 帧图片目录: {frames_dir}")
print(f" 请使用宿主 AI 的视觉能力逐帧查看并填写评分")
print(f" 评分结果写入: {scores_path}")
return data
# API 模式
api_key = os.environ.get("VIDEO_ANALYZER_API_KEY", "")
base_url = os.environ.get("VIDEO_ANALYZER_BASE_URL", "https://generativelanguage.googleapis.com/v1beta/openai")
model = os.environ.get("VIDEO_ANALYZER_MODEL", "gemini-2.0-flash")
if not api_key:
print("\n❌ API 模式需要设置环境变量 VIDEO_ANALYZER_API_KEY")
print(" export VIDEO_ANALYZER_API_KEY=\"your-api-key\"")
print(" export VIDEO_ANALYZER_BASE_URL=\"https://...\" # 可选")
print(" export VIDEO_ANALYZER_MODEL=\"gemini-2.0-flash\" # 可选")
sys.exit(1)
print(f"\n🤖 API 模式:使用 {model} 分析 {total_scenes} 个场景...")
print(f" API: {base_url}")
success_count = 0
for scene in scenes:
scene_num = scene.get("scene_number", 0)
frame_name = scene.get("filename", "").replace(".mp4", ".jpg")
frame_path = frames_dir / frame_name
if not frame_path.exists():
# 尝试其他命名
alt_name = f"{data.get('video_id', '')}-Scene-{scene_num:03d}.jpg"
frame_path = frames_dir / alt_name
if not frame_path.exists():
print(f" Scene {scene_num:03d}: ⚠️ 未找到帧图片,跳过")
continue
analysis = call_vision_api(
frame_path=frame_path,
scene_num=scene_num,
video_title=video_title,
total_scenes=total_scenes,
transcript_text=transcript_text[:500] if transcript_text else "",
api_key=api_key,
base_url=base_url,
model=model,
)
if analysis:
analysis = compute_weighted_score(analysis)
scene.update(analysis)
success_count += 1
print(f" Scene {scene_num:03d}: {analysis['selection']} | 加权 {analysis['weighted_score']:.2f} | {analysis.get('type_classification', 'N/A')}")
else:
print(f" Scene {scene_num:03d}: ❌ 分析失败")
# 限流:每次请求间隔
time.sleep(1.0)
# 保存
with open(scores_path, "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
print(f"\n✅ 评分完成:{success_count}/{total_scenes} 个场景成功")
print(f" 已保存到: {scores_path}")
return data
# ============================================================
# 精选镜头筛选与复制
# ============================================================
def select_and_copy_best_shots(scores_path: Path, threshold: float = 7.0) -> List[Dict]:
"""选择最佳镜头并复制到 best_shots 目录"""
with open(scores_path, "r", encoding="utf-8") as f:
data = json.load(f)
scenes = data.get("scenes", [])
video_dir = scores_path.parent
best_shots_dir = video_dir / "scenes" / "best_shots"
best_shots_dir.mkdir(exist_ok=True)
# 清空旧的精选
for old in best_shots_dir.glob("*.mp4"):
old.unlink()
# 筛选
best_shots = [
s for s in scenes
if s.get("weighted_score", 0) >= threshold or "MUST KEEP" in s.get("selection", "")
]
best_shots.sort(key=lambda x: x.get("weighted_score", 0), reverse=True)
print(f"\n⭐ 发现 {len(best_shots)} 个精选镜头 (阈值: {threshold})")
copied = []
for i, scene in enumerate(best_shots, 1):
src_path = Path(scene.get("file_path", ""))
if src_path.exists():
tag = scene.get("selection", "").replace("[", "").replace("]", "").replace(" ", "_")
dst_name = f"{i:02d}_{tag}_{src_path.name}"
dst_path = best_shots_dir / dst_name
shutil.copy2(src_path, dst_path)
copied.append(scene)
desc = scene.get("description", "N/A")[:30]
print(f" {i}. Scene {scene.get('scene_number', 0):03d} | {scene.get('weighted_score', 0):.2f} | {desc}...")
# 生成 README
_generate_readme(best_shots_dir, copied, data.get("video_id", "unknown"))
print(f"\n✅ 已复制 {len(copied)} 个精选镜头到: {best_shots_dir}")
return copied
def _generate_readme(best_shots_dir: Path, best_shots: List[Dict], video_id: str):
content = f"""# ⭐ 精选镜头 (Best Shots)
**视频 ID**: {video_id}
**入选数量**: {len(best_shots)} 个
**生成时间**: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}
## 精选列表
| 排名 | 场景 | 加权分 | 类型 | 描述 |
|------|------|--------|------|------|
"""
for i, s in enumerate(best_shots, 1):
content += f"| {i} | Scene {s.get('scene_number', 0):03d} | {s.get('weighted_score', 0):.2f} | {s.get('type_classification', 'N/A')} | {s.get('description', '')[:40]} |\n"
content += f"\n---\n*由 Video Expert Analyzer v2.0 筛选*\n"
(best_shots_dir / "README.md").write_text(content, encoding="utf-8")
# ============================================================
# 分析报告生成
# ============================================================
def generate_complete_report(scores_path: Path) -> Path:
with open(scores_path, "r", encoding="utf-8") as f:
data = json.load(f)
video_id = data.get("video_id", "unknown")
url = data.get("url", "")
scenes = data.get("scenes", [])
total = len(scenes)
if total == 0:
print("⚠️ 没有场景数据")
return scores_path
# 统计
scored_scenes = [s for s in scenes if "weighted_score" in s]
if not scored_scenes:
print("⚠️ 没有已评分的场景")
return scores_path
must_keep = sum(1 for s in scored_scenes if "MUST KEEP" in s.get("selection", ""))
usable = sum(1 for s in scored_scenes if "USABLE" in s.get("selection", ""))
discard = sum(1 for s in scored_scenes if "DISCARD" in s.get("selection", ""))
avg = sum(s["weighted_score"] for s in scored_scenes) / len(scored_scenes)
# 各维度平均
dims = ["aesthetic_beauty", "credibility", "impact", "memorability", "fun_interest"]
dim_avgs = {}
for d in dims:
vals = [s.get("scores", {}).get(d, 0) for s in scored_scenes if s.get("scores", {}).get(d)]
dim_avgs[d] = sum(vals) / len(vals) if vals else 0
report_path = scores_path.parent / f"{video_id}_complete_analysis.md"
sorted_scenes = sorted(scored_scenes, key=lambda x: x.get("weighted_score", 0), reverse=True)
# 构建报告
report = f"""# 🎬 视频专家分析报告 (Video Expert Analysis Report)
## 📋 基本信息
| 项目 | 内容 |
|------|------|
| **视频 ID** | {video_id} |
| **来源 URL** | {url} |
| **分析时间** | {datetime.now().strftime('%Y-%m-%d %H:%M:%S')} |
| **总场景数** | {total} |
| **已评分** | {len(scored_scenes)} |
| **平均加权得分** | {avg:.2f} |
### 筛选统计
| 等级 | 数量 | 占比 |
|------|------|------|
| 🌟 MUST KEEP | {must_keep} | {must_keep/len(scored_scenes)*100:.1f}% |
| 📁 USABLE | {usable} | {usable/len(scored_scenes)*100:.1f}% |
| 🗑️ DISCARD | {discard} | {discard/len(scored_scenes)*100:.1f}% |
### 各维度平均分
| 维度 | 平均分 |
|------|--------|
"""
for d in dims:
icon = "🟢" if dim_avgs[d] >= 7 else "🟡" if dim_avgs[d] >= 5 else "🔴"
report += f"| {get_term_chinese(d)} | {dim_avgs[d]:.2f} {icon} |\n"
report += f"""
---
## 🎞 场景排名
| 排名 | 场景 | 加权分 | 类型 | 等级 | 描述 |
|------|------|--------|------|------|------|
"""
for i, s in enumerate(sorted_scenes, 1):
desc = s.get("description", "N/A")[:30]
report += f"| {i} | Scene {s.get('scene_number', 0):03d} | **{s.get('weighted_score', 0):.2f}** | {s.get('type_classification', 'N/A')} | {s.get('selection', '')} | {desc} |\n"
report += f"""
---
## 📊 整体评价
### 综合评分: {avg:.2f} / 10
"""
if avg >= 8:
report += "🌟 **优秀** - 高质量素材,强烈推荐保留\n"
elif avg >= 6.5:
report += "📁 **良好** - 有可用价值,需要适当剪辑\n"
else:
report += "🗑️ **一般** - 整体质量较低\n"
report += f"""
---
*本报告由 Video Expert Analyzer v2.0 自动生成*
*分析时间: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}*
"""
report_path.write_text(report, encoding="utf-8")
print(f"✅ 完整分析报告已生成: {report_path}")
return report_path
# ============================================================
# CLI 入口
# ============================================================
if __name__ == "__main__":
if len(sys.argv) < 2:
print("""用法: python3 ai_analyzer.py <scene_scores.json> [--mode api|agent]
评分模式:
--mode api 通过远程 API 调用视觉大模型评分(需设置 VIDEO_ANALYZER_API_KEY)
--mode agent 生成评分模板,由宿主 AI 助手(如 IDE 中的 Gemini/Kimi)直接看图评分
环境变量 (API 模式):
VIDEO_ANALYZER_API_KEY API 密钥(必需)
VIDEO_ANALYZER_BASE_URL API 端点(默认 Gemini)
VIDEO_ANALYZER_MODEL 模型名称(默认 gemini-2.0-flash)
""")
sys.exit(1)
scores_path = Path(sys.argv[1])
video_dir = scores_path.parent
# 解析模式
mode = "api"
if "--mode" in sys.argv:
idx = sys.argv.index("--mode")
if idx + 1 < len(sys.argv):
mode = sys.argv[idx + 1]
print("=" * 60)
print(f"🤖 AI Scene Analyzer v2.0 ({mode.upper()} 模式)")
print("=" * 60)
# 1. 自动评分
data = auto_score_scenes(scores_path, video_dir, mode=mode)
if mode == "agent":
print("\n" + "=" * 60)
print("📝 Agent 模式说明")
print("=" * 60)
print(f"\n请使用宿主 AI 助手的视觉能力完成以下步骤:")
print(f" 1. 查看 {video_dir}/frames/ 中的每张截帧")
print(f" 2. 按 Walter Murch 法则五维度打分")
print(f" 3. 将结果更新到 {scores_path}")
print(f" 4. 再次运行本脚本(不带 --mode agent)生成报告")
sys.exit(0)
# 2. 复制精选镜头
print("\n" + "=" * 60)
print("⭐ 选择并复制精选镜头")
print("=" * 60)
select_and_copy_best_shots(scores_path, threshold=7.0)
# 3. 生成完整报告
print("\n" + "=" * 60)
print("📄 生成完整分析报告")
print("=" * 60)
report_path = generate_complete_report(scores_path)
print("\n" + "=" * 60)
print("✅ AI 分析完成!")
print("=" * 60)
print(f"\n📊 评分文件: {scores_path}")
print(f"⭐ 精选镜头: {video_dir}/scenes/best_shots/")
print(f"📄 完整报告: {report_path}")
#!/usr/bin/env python3
"""
Video Expert Analyzer v2.1 环境检测和依赖安装脚本
检测所有必要和可选依赖
"""
import subprocess
import sys
import shutil
def check_command(cmd: str, version_arg: str = "--version") -> tuple:
"""检查命令行工具是否可用"""
try:
result = subprocess.run(
[cmd, version_arg],
capture_output=True,
text=True,
timeout=10
)
version = result.stdout.strip() or result.stderr.strip()
return True, version.split('\n')[0]
except (FileNotFoundError, subprocess.TimeoutExpired):
return False, ""
def check_python_package(package: str) -> bool:
"""检查 Python 包是否已安装"""
try:
__import__(package)
return True
except ImportError:
return False
def main():
print("=" * 55)
print("🔍 Video Expert Analyzer v2.1 环境检测")
print("=" * 55)
print()
all_ok = True
missing_cmds = []
missing_pips = []
# ── 1. 系统工具 ──
print("1️⃣ 系统工具")
# 检查 FFmpeg
ok, version = check_command("ffmpeg", "-version")
if ok:
print(f" ✅ ffmpeg: {version[:60]}")
else:
print(f" ❌ ffmpeg 未安装 → brew install ffmpeg / 下载 FFmpeg")
all_ok = False
# 检查 yt-dlp (支持多种调用方式)
ok, version = False, ""
# 方法1: 直接命令
try:
result = subprocess.run(["yt-dlp", "--version"], capture_output=True, text=True, timeout=10)
if result.returncode == 0:
ok, version = True, result.stdout.strip()
except:
pass
# 方法2: 通过 py -m 调用
if not ok:
try:
result = subprocess.run([sys.executable, "-m", "yt_dlp", "--version"], capture_output=True, text=True, timeout=10)
if result.returncode == 0:
ok, version = True, result.stdout.strip()
except:
pass
if ok:
print(f" ✅ yt-dlp: {version}")
else:
print(f" ❌ yt-dlp 未安装 → pip3 install yt-dlp")
all_ok = False
# ── 2. 核心 Python 依赖(必需) ──
print("\n2️⃣ 核心 Python 依赖(必需)")
core_packages = {
"scenedetect": "scenedetect[opencv]",
"requests": "requests",
"torch": "torch",
}
for import_name, pip_name in core_packages.items():
if check_python_package(import_name):
print(f" ✅ {pip_name}")
else:
print(f" ❌ {pip_name} 未安装")
missing_pips.append(pip_name)
all_ok = False
# ── 3. 语音转录依赖(FunASR) ──
print("\n3️⃣ 语音转录 (FunASR)")
funasr_packages = {
"funasr": "funasr",
"modelscope": "modelscope",
"torchaudio": "torchaudio",
}
for import_name, pip_name in funasr_packages.items():
if check_python_package(import_name):
print(f" ✅ {pip_name}")
else:
print(f" ❌ {pip_name} 未安装")
missing_pips.append(pip_name)
all_ok = False
# ── 4. 可选依赖 ──
print("\n4️⃣ 可选依赖")
optional = {
"openai": ("openai", "API 模式评分"),
"rapidocr_onnxruntime": ("rapidocr-onnxruntime", "烧录字幕 OCR 检测"),
}
for import_name, (pip_name, desc) in optional.items():
if check_python_package(import_name):
print(f" ✅ {pip_name} ({desc})")
else:
print(f" ⚠️ {pip_name} 未安装 ({desc}) → pip3 install {pip_name}")
# ── 5. CUDA / MPS 检测 ──
print("\n5️⃣ GPU 加速")
try:
import torch
if torch.cuda.is_available():
print(f" ✅ CUDA: {torch.cuda.get_device_name(0)}")
elif hasattr(torch.backends, 'mps') and torch.backends.mps.is_available():
print(f" ✅ Apple MPS (Metal) 可用")
else:
print(" ⚠️ 无 GPU 加速,将使用 CPU(FunASR 转录可能较慢)")
except ImportError:
print(" ⚠️ PyTorch 未安装,无法检测 GPU")
# ── 结果汇总 ──
print()
print("=" * 55)
if all_ok:
print("✅ 所有核心依赖已满足!可以开始使用 Video Expert Analyzer。")
print()
print("快速开始:")
print(" python3 scripts/pipeline_enhanced.py --setup")
print(" python3 scripts/pipeline_enhanced.py <视频URL>")
else:
print("❌ 存在缺失依赖,请执行以下命令安装:")
print()
if missing_pips:
print(f" pip3 install {' '.join(missing_pips)}")
for cmd in missing_cmds:
print(f" {cmd}")
print()
print("或一键安装所有依赖:")
print(" pip3 install -r requirements.txt")
print("=" * 55)
return 0 if all_ok else 1
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""
抖音视频下载脚本
支持从抖音分享链接提取并下载视频(无水印版本)
使用方法:
python download_douyin.py <抖音链接> <输出路径>
示例:
python download_douyin.py "https://v.douyin.com/xxxxx" ./video.mp4
python download_douyin.py "https://www.douyin.com/video/xxxxx" ./video.mp4
"""
import requests
import re
import json
import sys
import os
from urllib.parse import unquote, urlparse
def is_douyin_url(url: str) -> bool:
"""检查是否为抖音链接"""
douyin_patterns = [
r'v\.douyin\.com',
r'www\.douyin\.com',
r'm\.douyin\.com',
r'douyin\.com/video/',
r'douyin\.com/jingxuan',
]
return any(re.search(pattern, url) for pattern in douyin_patterns)
def extract_video_id(url: str) -> str:
"""从抖音链接中提取视频ID"""
# 尝试从各种格式的链接中提取ID
patterns = [
r'/video/(\d+)',
r'modal_id=(\d+)',
r'share/video/(\d+)',
]
for pattern in patterns:
match = re.search(pattern, url)
if match:
return match.group(1)
# 如果是短链接,返回None,需要获取重定向后的URL
return None
def get_redirect_url(short_url: str) -> tuple:
"""获取重定向后的完整URL"""
headers = {
'User-Agent': 'Mozilla/5.0 (iPhone; CPU iPhone OS 13_2_3 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/13.0.3 Mobile/15E148 Safari/604.1',
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
'Accept-Language': 'zh-CN,zh;q=0.9',
}
try:
response = requests.get(short_url, headers=headers, allow_redirects=True, timeout=10)
return response.url, headers['User-Agent'], response.text
except Exception as e:
print(f"✗ 获取重定向URL失败: {e}")
return None, None, None
def extract_render_data(html: str) -> dict:
"""从HTML中提取RENDER_DATA"""
# 尝试多种可能的模式
patterns = [
r'<script id="RENDER_DATA" type="application/json">([^<]+)</script>',
r'window\._ROUTER_DATA\s*=\s*(\{.+?\});?\s*</script>',
r'window\._SSR_DATA\s*=\s*(\{.+?\});?\s*</script>',
r'window\._SSR_HYDRATED_DATA\s*=\s*(\{.+?\});?\s*</script>',
]
for pattern in patterns:
matches = re.findall(pattern, html, re.DOTALL)
if matches:
data_str = matches[0]
# URL解码
if '%' in data_str:
data_str = unquote(data_str)
try:
return json.loads(data_str)
except json.JSONDecodeError:
continue
return None
def extract_video_url(data: dict) -> str:
"""从RENDER_DATA中提取视频URL"""
def get_nested(obj, path):
"""安全地获取嵌套字典/列表值"""
current = obj
for key in path:
if isinstance(current, dict) and key in current:
current = current[key]
elif isinstance(current, list) and isinstance(key, int) and key < len(current):
current = current[key]
else:
return None
return current
# 尝试多种可能的路径
possible_paths = [
['loaderData', 'video_(id)/page', 'videoInfoRes', 'item_list', 0, 'video', 'play_addr', 'url_list'],
['loaderData', 'video_(id)/page', 'aweme_detail', 'video', 'play_addr', 'url_list'],
['videoInfoRes', 'item_list', 0, 'video', 'play_addr', 'url_list'],
['app', 'videoInfoRes', 'item_list', 0, 'video', 'play_addr', 'url_list'],
['app', 'videoDetail', 'video', 'play_addr', 'url_list'],
['video', 'play_addr', 'url_list'],
['aweme_detail', 'video', 'play_addr', 'url_list'],
]
for path in possible_paths:
url_list = get_nested(data, path)
if url_list and isinstance(url_list, list) and len(url_list) > 0:
video_url = url_list[0]
# 替换playwm为play获取无水印版本
video_url = video_url.replace('playwm', 'play')
return video_url
# 如果路径查找失败,尝试正则搜索
json_str = json.dumps(data)
play_patterns = [
r'"play_addr":\s*\{[^}]*"url_list":\s*\["([^"]+)"',
r'"playAddr":\s*\["([^"]+)"',
r'"download_addr":\s*\{[^}]*"url_list":\s*\["([^"]+)"',
]
for pattern in play_patterns:
matches = re.findall(pattern, json_str)
if matches:
video_url = matches[0].replace('playwm', 'play')
return video_url
return None
def download_video(video_url: str, output_path: str, user_agent: str) -> bool:
"""下载视频"""
headers = {
'User-Agent': user_agent,
'Referer': 'https://www.douyin.com/',
'Accept': '*/*',
'Accept-Language': 'zh-CN,zh;q=0.9',
}
try:
response = requests.get(video_url, headers=headers, stream=True, timeout=60)
if response.status_code not in [200, 206]:
print(f"✗ 下载失败,状态码: {response.status_code}")
return False
total_size = int(response.headers.get('content-length', 0))
downloaded = 0
with open(output_path, 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
downloaded += len(chunk)
if total_size > 0:
percent = (downloaded / total_size) * 100
print(f"\r进度: {percent:.1f}% ({downloaded:,}/{total_size:,} bytes)", end='', flush=True)
print() # 换行
return True
except Exception as e:
print(f"✗ 下载视频时出错: {e}")
return False
def download_douyin_video(url: str, output_path: str) -> bool:
"""
下载抖音视频的主函数
Args:
url: 抖音视频链接(支持短链接和长链接)
output_path: 输出文件路径
Returns:
bool: 下载是否成功
"""
print(f"🎬 开始下载抖音视频")
print(f" 链接: {url}")
print(f" 输出: {output_path}")
print()
# 步骤1: 获取重定向URL和页面内容
print("步骤 1/4: 获取页面信息...")
full_url, user_agent, html = get_redirect_url(url)
if not full_url:
return False
print(f"✓ 获取到页面 ({len(html):,} 字符)")
# 步骤2: 提取RENDER_DATA
print("\n步骤 2/4: 提取视频数据...")
render_data = extract_render_data(html)
if not render_data:
print("✗ 无法提取视频数据")
return False
print("✓ 提取到视频数据")
# 步骤3: 提取视频URL
print("\n步骤 3/4: 解析视频地址...")
video_url = extract_video_url(render_data)
if not video_url:
print("✗ 无法获取视频下载地址")
return False
print(f"✓ 获取到视频地址")
# 步骤4: 下载视频
print("\n步骤 4/4: 下载视频...")
success = download_video(video_url, output_path, user_agent)
if success:
file_size = os.path.getsize(output_path)
print(f"✓ 下载完成: {file_size:,} bytes")
return True
else:
return False
def main():
if len(sys.argv) < 3:
print("用法: python download_douyin.py <抖音链接> <输出路径>")
print("示例: python download_douyin.py 'https://v.douyin.com/xxxxx' ./video.mp4")
sys.exit(1)
url = sys.argv[1]
output_path = sys.argv[2]
# 检查是否为抖音链接
if not is_douyin_url(url):
print(f"✗ 不是有效的抖音链接: {url}")
sys.exit(1)
success = download_douyin_video(url, output_path)
sys.exit(0 if success else 1)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
智能字幕提取脚本 - FunASR + RapidOCR 版本
流程:B站API字幕 → 内嵌字幕 → 烧录字幕检测(RapidOCR) → FunASR语音转录
技术栈:
- B站 API: 直接获取平台字幕(需 cookies)
- RapidOCR (ONNX): 轻量级 OCR,用于提取烧录字幕
- FunASR: 中文语音转录,配合 VAD 分段和标点模型
"""
import subprocess
import sys
import os
import re
import tempfile
from pathlib import Path
import json
# ============================================================
# L0: B站 API 字幕获取(最高优先级)
# ============================================================
def extract_bvid(video_url_or_path: str) -> str:
"""从 URL 或文件名中提取 B站 BV 号"""
# 匹配 BV 号模式(BV + 10位字母数字)
match = re.search(r'(BV[a-zA-Z0-9]{10})', video_url_or_path)
if match:
return match.group(1)
return ""
def get_bilibili_subtitle(bvid: str, output_srt: str) -> bool:
"""
通过 B站 API 获取字幕
自动从浏览器读取 cookies,无需手动配置
优先级: yt-dlp cookies > browser_cookie3 > 配置文件/环境变量
"""
# 调用独立的字幕获取脚本
script_dir = os.path.dirname(os.path.abspath(__file__))
fetch_script = os.path.join(script_dir, "fetch_bilibili_subtitle.py")
if os.path.exists(fetch_script):
try:
cmd = [sys.executable, fetch_script, bvid, output_srt]
result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
print(result.stdout)
if result.stderr:
print(result.stderr)
if result.returncode == 0 and os.path.exists(output_srt):
# 检查文件是否有实际内容
if os.path.getsize(output_srt) > 10:
return True
return False
except subprocess.TimeoutExpired:
print(" ⚠️ 字幕获取超时")
return False
except Exception as e:
print(f" ⚠️ 调用字幕获取脚本失败: {e}")
return False
else:
print(f" ⚠️ 未找到 fetch_bilibili_subtitle.py 脚本")
# 回退到简单的无 cookies 尝试
return _simple_bilibili_fetch(bvid, output_srt)
def _simple_bilibili_fetch(bvid: str, output_srt: str) -> bool:
"""简单的 B站字幕获取(无 cookies,通常会失败但不影响流程)"""
try:
import requests
except ImportError:
return False
headers = {
"User-Agent": "Mozilla/5.0",
"Referer": "https://www.bilibili.com",
}
try:
resp = requests.get(
f"https://api.bilibili.com/x/player/pagelist?bvid={bvid}",
headers=headers, timeout=10
)
data = resp.json()
if data.get("code") != 0 or not data.get("data"):
return False
cid = data["data"][0]["cid"]
resp = requests.get(
f"https://api.bilibili.com/x/web-interface/view?bvid={bvid}",
headers=headers, timeout=10
)
aid = resp.json()["data"]["aid"]
resp = requests.get(
f"https://api.bilibili.com/x/player/wbi/v2?aid={aid}&cid={cid}",
headers=headers, timeout=10
)
subtitles = resp.json().get("data", {}).get("subtitle", {}).get("subtitles", [])
if not subtitles:
return False
sub_url = subtitles[0].get("subtitle_url", "")
if sub_url.startswith("//"):
sub_url = "https:" + sub_url
resp = requests.get(sub_url, headers=headers, timeout=10)
body = resp.json().get("body", [])
if not body:
return False
with open(output_srt, 'w', encoding='utf-8') as f:
for i, item in enumerate(body, 1):
start = format_timestamp(item.get("from", 0))
end = format_timestamp(item.get("to", 0))
content = item.get("content", "").strip()
if content:
f.write(f"{i}\n{start} --> {end}\n{content}\n\n")
return True
except Exception:
return False
# ============================================================
# L1: 内嵌字幕检测
# ============================================================
def check_embedded_subtitle(video_path: str) -> tuple[bool, str]:
"""
检查视频是否包含内嵌字幕流
返回: (是否有内嵌字幕, 字幕文件路径或错误信息)
"""
try:
cmd = [
"ffprobe", "-v", "quiet", "-print_format", "json",
"-show_streams", video_path
]
result = subprocess.run(cmd, capture_output=True, text=True, check=True)
data = json.loads(result.stdout)
streams = data.get("streams", [])
subtitle_streams = [s for s in streams if s.get("codec_type") == "subtitle"]
if subtitle_streams:
output_srt = video_path.rsplit(".", 1)[0] + "_embedded.srt"
cmd = [
"ffmpeg", "-y", "-i", video_path,
"-map", f"0:s:0", output_srt
]
subprocess.run(cmd, capture_output=True, check=True)
return True, output_srt
else:
return False, "无内嵌字幕流"
except Exception as e:
return False, f"检测失败: {e}"
# ============================================================
# L2: 烧录字幕检测与提取 (RapidOCR)
# ============================================================
def capture_frame(video_path: str, timestamp: str = "00:00:05") -> str:
"""截取视频指定时间的帧"""
try:
with tempfile.NamedTemporaryFile(suffix=".jpg", delete=False) as tmp:
frame_path = tmp.name
cmd = [
"ffmpeg", "-y", "-ss", timestamp, "-i", video_path,
"-vframes", "1", "-q:v", "2", frame_path
]
subprocess.run(cmd, capture_output=True, check=True)
return frame_path
except Exception as e:
return ""
def _format_time_hms(seconds: int) -> str:
"""将秒数格式化为 HH:MM:SS 格式(用于 ffmpeg 时间戳)"""
h = seconds // 3600
m = (seconds % 3600) // 60
s = seconds % 60
return f"{h:02d}:{m:02d}:{s:02d}"
def check_burned_subtitle(frame_path: str) -> bool:
"""使用 RapidOCR 检测画面是否有烧录字幕"""
try:
from rapidocr_onnxruntime import RapidOCR
ocr = RapidOCR()
result = ocr(frame_path)
# 如果检测到文字,认为有烧录字幕
if result and result[0]:
text_count = len([line for line in result[0] if line])
# 检测到至少2行文字,认为是字幕
return text_count >= 2
return False
except ImportError:
print("⚠️ RapidOCR 未安装,跳过烧录字幕检测")
print(" 安装命令: pip install rapidocr-onnxruntime")
return False
except Exception as e:
print(f"⚠️ OCR 检测失败: {e}")
return False
def extract_burned_subtitle_ocr(video_path: str, output_srt: str) -> bool:
"""使用 RapidOCR 提取烧录字幕"""
try:
from rapidocr_onnxruntime import RapidOCR
print("🔍 使用 RapidOCR 提取烧录字幕...")
cmd = [
"ffprobe", "-v", "error", "-show_entries",
"format=duration", "-of", "default=noprint_wrappers=1:nokey=1",
video_path
]
result = subprocess.run(cmd, capture_output=True, text=True)
duration = float(result.stdout.strip())
ocr = RapidOCR()
# 每隔2秒截取一帧进行 OCR(减少计算量)
subtitles = []
for t in range(0, int(duration), 2):
timestamp = _format_time_hms(t)
frame_path = capture_frame(video_path, timestamp)
if not frame_path:
continue
result = ocr(frame_path)
if result and result[0]:
# 提取文字
texts = []
for line in result[0]:
if line:
text = line[1]
confidence = line[2]
# 修复: confidence 可能是 str 类型,统一转为 float
try:
conf = float(confidence)
except (ValueError, TypeError):
conf = 0.0
if conf > 0.7: # 置信度阈值
texts.append(text)
if texts:
start_ts = format_timestamp(t)
end_ts = format_timestamp(t + 2)
subtitles.append({
'index': len(subtitles) + 1,
'start': start_ts,
'end': end_ts,
'text': ' '.join(texts)
})
os.unlink(frame_path)
# 写入 SRT 文件
with open(output_srt, 'w', encoding='utf-8') as f:
for sub in subtitles:
f.write(f"{sub['index']}\n")
f.write(f"{sub['start']} --> {sub['end']}\n")
f.write(f"{sub['text']}\n\n")
print(f"✅ OCR 提取完成: {len(subtitles)} 条字幕")
return True
except Exception as e:
print(f"❌ OCR 提取失败: {e}")
return False
# ============================================================
# L3: FunASR 语音转录
# ============================================================
def extract_audio(video_path: str, audio_path: str) -> bool:
"""从视频中提取音频"""
try:
cmd = [
"ffmpeg", "-y", "-i", video_path,
"-vn", "-acodec", "pcm_s16le",
"-ar", "16000", "-ac", "1",
audio_path
]
subprocess.run(cmd, capture_output=True, check=True)
return True
except subprocess.CalledProcessError as e:
print(f"❌ 音频提取失败: {e}")
return False
def format_timestamp(seconds: float) -> str:
"""格式化时间戳为 SRT 格式"""
hours = int(seconds // 3600)
minutes = int((seconds % 3600) // 60)
secs = int(seconds % 60)
millis = int((seconds % 1) * 1000)
return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
def _split_text_by_punctuation(text: str, timestamps: list) -> list:
"""
按标点符号切分带字级时间戳的文本为自然句
timestamps: [[start_ms, end_ms], ...] 每个字/词的时间戳
返回: [{'text': str, 'start_ms': int, 'end_ms': int}, ...]
"""
# 句末标点
sentence_endings = set('。!?!?;;…')
# 次级切分标点(逗号等,仅在句子过长时切)
clause_breaks = set(',,、')
sentences = []
current_chars = []
current_start_idx = 0
ts_len = len(timestamps)
text_len = len(text)
for char_idx, char in enumerate(text):
current_chars.append(char)
# 映射字符位置到时间戳位置
ts_idx = min(int(char_idx / text_len * ts_len), ts_len - 1) if ts_len > 0 else 0
is_end = char in sentence_endings
is_clause = char in clause_breaks and len(current_chars) > 25 # 逗号切分仅在 >25 字时
is_last = char_idx == text_len - 1
if is_end or is_clause or is_last:
sent_text = ''.join(current_chars).strip()
if sent_text:
start_ts_idx = min(int(current_start_idx / text_len * ts_len), ts_len - 1) if ts_len > 0 else 0
end_ts_idx = ts_idx
start_ms = timestamps[start_ts_idx][0] if ts_len > 0 else 0
end_ms = timestamps[end_ts_idx][1] if ts_len > 0 else 0
sentences.append({
'text': sent_text,
'start_ms': start_ms,
'end_ms': end_ms,
})
current_chars = []
current_start_idx = char_idx + 1
return sentences
def extract_with_funasr(video_path: str, output_srt: str) -> bool:
"""
使用 FunASR 进行语音转录
配合 VAD 分段模型 + 标点模型,正确处理长音频
调用方式参照 FunASR 官方 demo:
https://github.com/modelscope/FunASR/blob/main/examples/industrial_data_pretraining/paraformer/demo.py
"""
try:
from funasr import AutoModel
print("🎤 使用 FunASR 进行语音转录...")
print(" ASR 模型: paraformer-zh (含 VAD + 标点)")
print(" ⚠️ 首次运行需下载约 2-3GB 模型文件,请耐心等待")
# 提取音频
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:
audio_path = tmp.name
if not extract_audio(video_path, audio_path):
return False
# 加载 FunASR 模型(官方推荐的短名称 + VAD + 标点)
model = AutoModel(
model="paraformer-zh",
vad_model="fsmn-vad",
vad_kwargs={"max_single_segment_time": 60000},
punc_model="ct-punc",
disable_update=True,
)
# 转录(VAD 自动分段,标点自动恢复,cache={} 是官方推荐参数)
result = model.generate(
input=audio_path,
batch_size_s=300,
cache={},
)
# 生成 SRT
subtitle_count = 0
with open(output_srt, 'w', encoding='utf-8') as f:
for res in result:
text = res.get('text', '').strip()
timestamps = res.get('timestamp', [])
sentence_info = res.get('sentence_info', [])
if sentence_info:
# 方案A: 使用句级时间戳(最佳,如果模型返回了)
for sent in sentence_info:
sent_text = sent.get('text', '').strip()
if sent_text:
subtitle_count += 1
start = format_timestamp(sent.get('start', 0) / 1000)
end = format_timestamp(sent.get('end', 0) / 1000)
f.write(f"{subtitle_count}\n{start} --> {end}\n{sent_text}\n\n")
elif timestamps and text:
# 方案B: 按标点符号切分 + 字级时间戳映射
sentences = _split_text_by_punctuation(text, timestamps)
for sent in sentences:
subtitle_count += 1
start = format_timestamp(sent['start_ms'] / 1000)
end = format_timestamp(sent['end_ms'] / 1000)
f.write(f"{subtitle_count}\n{start} --> {end}\n{sent['text']}\n\n")
elif text:
# 方案C: 无时间戳,仅输出文本
subtitle_count += 1
f.write(f"{subtitle_count}\n00:00:00,000 --> 00:00:00,000\n{text}\n\n")
# 清理临时文件
os.unlink(audio_path)
print(f"✅ FunASR 转录完成: {subtitle_count} 条字幕")
return subtitle_count > 0
except ImportError:
print("❌ FunASR 未安装")
print(" 安装命令: pip install funasr modelscope torchaudio")
return False
except Exception as e:
print(f"❌ FunASR 转录失败: {e}")
import traceback
traceback.print_exc()
return False
# ============================================================
# 主流程
# ============================================================
def smart_subtitle_extraction(video_path: str, output_srt: str, video_url: str = "") -> tuple[bool, str]:
"""
智能字幕提取主函数
流程: B站API字幕 → 内嵌字幕 → 烧录字幕(RapidOCR) → FunASR语音转录
返回: (是否成功, 使用的模式)
"""
print("=" * 50)
print("🎬 智能字幕提取 (B站API + RapidOCR + FunASR)")
print("=" * 50)
print(f"视频: {video_path}")
print()
# 步骤0: 尝试从B站API获取字幕(最优先)
bvid = extract_bvid(video_url) or extract_bvid(video_path)
if bvid:
print("步骤 0/4: 尝试B站API字幕获取...")
if get_bilibili_subtitle(bvid, output_srt):
return True, "bilibili_api"
print()
# 步骤1: 检查内嵌字幕
print("步骤 1/3: 检查内嵌字幕...")
has_embedded, result = check_embedded_subtitle(video_path)
if has_embedded:
print(f"✅ 发现内嵌字幕,已提取: {result}")
if result != output_srt:
import shutil
shutil.copy(result, output_srt)
return True, "embedded"
else:
print(f"⚠️ {result}")
# 步骤2: 检测烧录字幕
print("\n步骤 2/3: 检测烧录字幕 (RapidOCR)...")
frame_path = capture_frame(video_path, "00:00:05")
if frame_path:
has_burned = check_burned_subtitle(frame_path)
os.unlink(frame_path)
if has_burned:
print("✅ 检测到烧录字幕,使用 RapidOCR 提取...")
if extract_burned_subtitle_ocr(video_path, output_srt):
return True, "ocr"
else:
print("⚠️ 未检测到烧录字幕")
# 步骤3: 使用 FunASR
print("\n步骤 3/3: 使用 FunASR 语音转录...")
if extract_with_funasr(video_path, output_srt):
return True, "funasr"
return False, "failed"
def main():
if len(sys.argv) < 3:
print("用法: python extract_subtitle_funasr.py <视频路径> <输出SRT路径> [视频URL]")
print()
print("参数说明:")
print(" 视频路径 - 本地视频文件路径")
print(" 输出SRT - 输出的 SRT 字幕文件路径")
print(" 视频URL - 可选,原始视频URL(用于B站API字幕获取)")
sys.exit(1)
video_path = sys.argv[1]
output_srt = sys.argv[2]
video_url = sys.argv[3] if len(sys.argv) > 3 else ""
if not os.path.exists(video_path):
print(f"❌ 视频文件不存在: {video_path}")
sys.exit(1)
success, mode = smart_subtitle_extraction(video_path, output_srt, video_url)
if success:
print(f"\n✅ 字幕提取成功!")
print(f" 模式: {mode}")
print(f" 输出: {output_srt}")
sys.exit(0)
else:
print(f"\n❌ 字幕提取失败")
sys.exit(1)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
B站字幕获取脚本 - 自动从浏览器读取 cookies 并获取视频字幕
功能:
- 自动检测已登录的浏览器并获取 B站 cookies
- 通过 B站 API 直接获取 AI 生成字幕(比本地 ASR 更快更准)
- 输出标准 SRT 格式字幕文件
Cookies 获取优先级:
1. yt-dlp --cookies-from-browser(最可靠,持续跟进浏览器加密更新)
2. browser_cookie3 Python 库(备选)
3. 手动配置 ~/.bilibili_cookies.txt 或环境变量(兜底)
用法:
python fetch_bilibili_subtitle.py <BV号或URL> <输出SRT路径> [--browser chrome|firefox|safari|edge]
示例:
python fetch_bilibili_subtitle.py BV1vdZ6BJEcQ output.srt
python fetch_bilibili_subtitle.py "https://www.bilibili.com/video/BV1vdZ6BJEcQ/" output.srt
python fetch_bilibili_subtitle.py BV1vdZ6BJEcQ output.srt --browser firefox
"""
import subprocess
import sys
import os
import re
import json
import tempfile
import argparse
from pathlib import Path
try:
import requests
except ImportError:
print("❌ requests 未安装: pip install requests")
sys.exit(1)
# ============================================================
# BV号/URL 解析
# ============================================================
def extract_bvid(input_str: str) -> str:
"""从 URL、BV号 或文件名中提取 BV号"""
# 直接是 BV号
match = re.search(r'(BV[a-zA-Z0-9]{10})', input_str)
if match:
return match.group(1)
# 短链需要解析重定向
if 'b23.tv' in input_str:
try:
resp = requests.head(input_str, allow_redirects=True, timeout=10)
match = re.search(r'(BV[a-zA-Z0-9]{10})', resp.url)
if match:
return match.group(1)
except Exception:
pass
return ""
# ============================================================
# Cookies 获取策略
# ============================================================
def get_cookies_via_ytdlp(browser: str = "chrome") -> dict:
"""
策略1: 通过 yt-dlp 从浏览器获取 cookies(最可靠)
yt-dlp 持续跟进浏览器加密更新,比第三方库更稳定
"""
print(f" 🔑 尝试从 {browser} 获取 cookies (via yt-dlp)...")
try:
# 用 yt-dlp 导出 cookies 到临时文件
with tempfile.NamedTemporaryFile(suffix=".txt", delete=False, mode='w') as tmp:
cookies_file = tmp.name
cmd = [
"yt-dlp",
"--cookies-from-browser", browser,
"--cookies", cookies_file,
"--skip-download",
"--no-warnings",
"-q",
"https://www.bilibili.com/video/BV1xx411c7mD/", # 任意有效BV号
]
result = subprocess.run(
cmd,
capture_output=True,
text=True,
timeout=30,
)
if os.path.exists(cookies_file) and os.path.getsize(cookies_file) > 0:
cookies = _parse_netscape_cookies(cookies_file, ".bilibili.com")
os.unlink(cookies_file)
if cookies.get("SESSDATA"):
print(f" ✅ 成功获取 cookies (SESSDATA={cookies['SESSDATA'][:8]}...)")
return cookies
else:
print(f" ⚠️ cookies 中无 SESSDATA(可能未登录B站)")
return {}
else:
os.unlink(cookies_file) if os.path.exists(cookies_file) else None
print(f" ⚠️ yt-dlp 未能导出 cookies")
return {}
except FileNotFoundError:
print(f" ⚠️ yt-dlp 未安装,跳过")
return {}
except subprocess.TimeoutExpired:
print(f" ⚠️ yt-dlp 超时(可能需要系统钥匙串权限)")
return {}
except Exception as e:
print(f" ⚠️ yt-dlp 获取失败: {e}")
return {}
def get_cookies_via_browser_cookie3(browser: str = "chrome") -> dict:
"""
策略2: 通过 browser_cookie3 库获取 cookies
注意: Chrome 2024年后新加密可能导致部分 cookies 值为空
"""
print(f" 🔑 尝试从 {browser} 获取 cookies (via browser_cookie3)...")
try:
import browser_cookie3
browser_map = {
"chrome": browser_cookie3.chrome,
"firefox": browser_cookie3.firefox,
"edge": browser_cookie3.edge,
"opera": browser_cookie3.opera,
}
if browser not in browser_map:
print(f" ⚠️ browser_cookie3 不支持 {browser}")
return {}
cj = browser_map[browser](domain_name=".bilibili.com")
cookies = {}
for cookie in cj:
if cookie.domain and ".bilibili.com" in cookie.domain:
cookies[cookie.name] = cookie.value
if cookies.get("SESSDATA"):
print(f" ✅ 成功获取 cookies (SESSDATA={cookies['SESSDATA'][:8]}...)")
return cookies
else:
print(f" ⚠️ cookies 中无 SESSDATA(可能未登录或加密问题)")
return {}
except ImportError:
print(f" ⚠️ browser_cookie3 未安装,跳过 (pip install browser_cookie3)")
return {}
except Exception as e:
print(f" ⚠️ browser_cookie3 获取失败: {e}")
return {}
def get_cookies_from_config() -> dict:
"""
策略3: 从配置文件或环境变量获取 cookies(兜底方案)
支持:
- 环境变量: BILIBILI_SESSDATA, BILIBILI_BILI_JCT
- cookies 文件: ~/.bilibili_cookies.txt (Netscape 格式)
"""
print(" 🔑 尝试从配置文件/环境变量获取 cookies...")
cookies = {}
# 方式A: 环境变量
sessdata = os.environ.get("BILIBILI_SESSDATA", "")
if sessdata:
cookies["SESSDATA"] = sessdata
bili_jct = os.environ.get("BILIBILI_BILI_JCT", "")
if bili_jct:
cookies["bili_jct"] = bili_jct
print(f" ✅ 从环境变量获取 (SESSDATA={sessdata[:8]}...)")
return cookies
# 方式B: Netscape cookies 文件
cookies_file = os.path.expanduser("~/.bilibili_cookies.txt")
if os.path.exists(cookies_file):
cookies = _parse_netscape_cookies(cookies_file, ".bilibili.com")
if cookies.get("SESSDATA"):
print(f" ✅ 从 {cookies_file} 获取 (SESSDATA={cookies['SESSDATA'][:8]}...)")
return cookies
print(" ⚠️ 未找到配置的 cookies")
return {}
def _parse_netscape_cookies(filepath: str, domain_filter: str = "") -> dict:
"""解析 Netscape 格式 cookies 文件"""
cookies = {}
try:
with open(filepath, 'r') as f:
for line in f:
line = line.strip()
if line and not line.startswith('#'):
parts = line.split('\t')
if len(parts) >= 7:
domain = parts[0]
name = parts[5]
value = parts[6]
if not domain_filter or domain_filter in domain:
cookies[name] = value
except Exception:
pass
return cookies
def get_bilibili_cookies(preferred_browser: str = "chrome") -> dict:
"""
按优先级尝试获取 B站 cookies
优先级:yt-dlp > browser_cookie3 > 配置文件/环境变量
"""
print("\n📦 获取 B站 cookies...")
# 要尝试的浏览器列表
browsers = [preferred_browser]
for b in ["chrome", "firefox", "edge", "safari"]:
if b not in browsers:
browsers.append(b)
# 策略1: yt-dlp(按浏览器优先级)
for browser in browsers:
cookies = get_cookies_via_ytdlp(browser)
if cookies.get("SESSDATA"):
return cookies
# 策略2: browser_cookie3(按浏览器优先级)
for browser in browsers:
cookies = get_cookies_via_browser_cookie3(browser)
if cookies.get("SESSDATA"):
return cookies
# 策略3: 配置文件/环境变量
cookies = get_cookies_from_config()
if cookies.get("SESSDATA"):
return cookies
return {}
# ============================================================
# B站 API 字幕获取
# ============================================================
def fetch_subtitle(bvid: str, cookies: dict, output_srt: str) -> bool:
"""通过 B站 API 获取字幕并保存为 SRT"""
headers = {
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36",
"Referer": "https://www.bilibili.com",
}
try:
# 步骤1: BV号 → CID
print("\n📡 调用 B站 API...")
url = f"https://api.bilibili.com/x/player/pagelist?bvid={bvid}"
resp = requests.get(url, headers=headers, cookies=cookies, timeout=10)
data = resp.json()
if data.get("code") != 0 or not data.get("data"):
print(f" ❌ 获取视频信息失败: {data.get('message', '未知错误')}")
return False
cid = data["data"][0]["cid"]
part_name = data["data"][0].get("part", "")
duration = data["data"][0].get("duration", 0)
print(f" 📺 视频: {part_name} (时长: {duration}s, CID: {cid})")
# 步骤2: 获取 AID
url = f"https://api.bilibili.com/x/web-interface/view?bvid={bvid}"
resp = requests.get(url, headers=headers, cookies=cookies, timeout=10)
data = resp.json()
if data.get("code") != 0:
print(f" ❌ 获取AID失败: {data.get('message', '未知错误')}")
return False
aid = data["data"]["aid"]
title = data["data"].get("title", "")
print(f" 📝 标题: {title}")
# 步骤3: 获取字幕列表
url = f"https://api.bilibili.com/x/player/wbi/v2?aid={aid}&cid={cid}"
resp = requests.get(url, headers=headers, cookies=cookies, timeout=10)
data = resp.json()
if data.get("code") != 0:
print(f" ❌ 获取字幕信息失败: {data.get('message', '未知错误')}")
return False
subtitles = data.get("data", {}).get("subtitle", {}).get("subtitles", [])
if not subtitles:
print(" ⚠️ 该视频无字幕(未开启AI字幕或需要登录)")
return False
# 显示可用字幕
print(f" 📋 可用字幕: {len(subtitles)} 条")
for s in subtitles:
print(f" - {s.get('lan_doc', '?')} ({s.get('lan', '?')})")
# 选择中文字幕(优先 ai-zh)
chosen = subtitles[0]
for s in subtitles:
if s.get("lan") in ("ai-zh", "zh-Hans", "zh-CN", "zh"):
chosen = s
break
subtitle_url = chosen.get("subtitle_url", "")
if not subtitle_url:
print(" ❌ 字幕URL为空")
return False
if subtitle_url.startswith("//"):
subtitle_url = "https:" + subtitle_url
# 步骤4: 下载字幕 JSON
resp = requests.get(subtitle_url, headers=headers, timeout=10)
subtitle_data = resp.json()
body = subtitle_data.get("body", [])
if not body:
print(" ❌ 字幕内容为空")
return False
# 步骤5: 转换为 SRT 格式
with open(output_srt, 'w', encoding='utf-8') as f:
for i, item in enumerate(body, 1):
start = item.get("from", 0)
end = item.get("to", 0)
content = item.get("content", "").strip()
if content:
start_ts = _format_srt_timestamp(start)
end_ts = _format_srt_timestamp(end)
f.write(f"{i}\n{start_ts} --> {end_ts}\n{content}\n\n")
print(f"\n✅ 字幕获取成功!")
print(f" 条数: {len(body)}")
print(f" 语言: {chosen.get('lan_doc', '未知')}")
print(f" 输出: {output_srt}")
return True
except requests.exceptions.RequestException as e:
print(f" ❌ 网络请求失败: {e}")
return False
except Exception as e:
print(f" ❌ 字幕获取失败: {e}")
import traceback
traceback.print_exc()
return False
def _format_srt_timestamp(seconds: float) -> str:
"""格式化时间戳为 SRT 格式 HH:MM:SS,mmm"""
hours = int(seconds // 3600)
minutes = int((seconds % 3600) // 60)
secs = int(seconds % 60)
millis = int((seconds % 1) * 1000)
return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
# ============================================================
# 主入口
# ============================================================
def main():
parser = argparse.ArgumentParser(
description="从 B站获取视频字幕(自动读取浏览器 cookies)",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
示例:
python fetch_bilibili_subtitle.py BV1vdZ6BJEcQ output.srt
python fetch_bilibili_subtitle.py "https://b23.tv/W2ot8As" output.srt
python fetch_bilibili_subtitle.py BV1vdZ6BJEcQ output.srt --browser firefox
Cookies 获取优先级:
1. yt-dlp --cookies-from-browser(最可靠)
2. browser_cookie3 Python 库
3. ~/.bilibili_cookies.txt 或环境变量 BILIBILI_SESSDATA
如果所有自动方式都失败,可以手动配置:
export BILIBILI_SESSDATA="你的SESSDATA值"
export BILIBILI_BILI_JCT="你的bili_jct值"
"""
)
parser.add_argument("input", help="B站 BV号 或 视频URL")
parser.add_argument("output", help="输出 SRT 文件路径")
parser.add_argument("--browser", default="chrome",
help="优先使用的浏览器 (default: chrome)")
args = parser.parse_args()
# 解析 BV号
print("=" * 50)
print("🎬 B站字幕获取工具")
print("=" * 50)
bvid = extract_bvid(args.input)
if not bvid:
print(f"❌ 无法解析 BV号: {args.input}")
sys.exit(1)
print(f"📌 BV号: {bvid}")
# 获取 cookies
cookies = get_bilibili_cookies(args.browser)
if not cookies.get("SESSDATA"):
print("\n❌ 无法获取 B站 cookies,请确保:")
print(" 1. 已在浏览器中登录 bilibili.com")
print(" 2. 已安装 yt-dlp: pip install yt-dlp")
print(" 3. 或安装 browser_cookie3: pip install browser_cookie3")
print(" 4. 或手动设置: export BILIBILI_SESSDATA='你的值'")
sys.exit(1)
# 获取字幕
success = fetch_subtitle(bvid, cookies, args.output)
if success:
sys.exit(0)
else:
print("\n💡 提示: 如果字幕获取失败,可以回退到本地转录方式:")
print(" python extract_subtitle_funasr.py <视频文件> <输出SRT>")
sys.exit(1)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
XiaoHongShu Video Downloader
小红书视频下载模块 - 用于处理小红书视频笔记
"""
import requests
import re
import json
import sys
from urllib.parse import unquote, urlparse
from pathlib import Path
from typing import Optional, Tuple, Dict
class XiaohongshuDownloader:
"""小红书视频下载器"""
def __init__(self):
# 使用移动端UA更容易获取数据
self.mobile_user_agent = 'Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.0 Mobile/15E148 Safari/604.1'
# PC端UA
self.pc_user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
self.session = requests.Session()
@staticmethod
def is_xiaohongshu_url(url: str) -> bool:
"""检查URL是否为小红书链接"""
xhs_patterns = [
'xiaohongshu.com',
'xhslink.com',
]
return any(pattern in url.lower() for pattern in xhs_patterns)
@staticmethod
def extract_note_id(url: str) -> Optional[str]:
"""从URL中提取笔记ID"""
# 处理 /discovery/item/xxxx 格式
match = re.search(r'/item/([a-f0-9]+)', url)
if match:
return match.group(1)
# 处理 /explore/xxxx 格式
match = re.search(r'/explore/([a-f0-9]+)', url)
if match:
return match.group(1)
# 处理短链接 xhslink.com/xxxx
if 'xhslink.com' in url:
return None # 短链接需要重定向获取
return None
def get_redirect_url(self, short_url: str) -> Tuple[Optional[str], str]:
"""获取短链接重定向后的完整URL"""
try:
self.session.headers.update({
'User-Agent': self.mobile_user_agent,
})
response = self.session.get(short_url, allow_redirects=True, timeout=10)
return response.url, self.mobile_user_agent
except Exception as e:
print(f" ⚠️ 获取重定向URL失败: {e}")
return None, self.mobile_user_agent
def get_note_info(self, url: str) -> Dict:
"""获取笔记信息(标题、作者、视频URL等)"""
result = {
"success": False,
"title": "",
"uploader": "",
"video_url": None,
"cover_url": None,
"note_id": None,
"platform": "xiaohongshu"
}
# 处理短链接
if 'xhslink.com' in url:
url, _ = self.get_redirect_url(url)
if not url:
return result
# 提取笔记ID
note_id = self.extract_note_id(url)
result["note_id"] = note_id
try:
# 使用PC UA获取页面(通常内容更完整)
self.session.headers.update({
'User-Agent': self.pc_user_agent,
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
'Accept-Language': 'zh-CN,zh;q=0.9,en;q=0.8',
'Referer': 'https://www.xiaohongshu.com/',
})
response = self.session.get(url, timeout=15)
html = response.text
# 尝试多种方式提取数据
# 方式1: 从 __INITIAL_STATE__ 提取
state_match = re.search(r'window\.__INITIAL_STATE__\s*=\s*(\{.+?\})\s*</script>', html, re.DOTALL)
if state_match:
try:
# 处理undefined等非标准JSON
state_str = state_match.group(1)
state_str = re.sub(r':undefined', ':null', state_str)
state_str = re.sub(r':NaN', ':null', state_str)
state = json.loads(state_str)
# 提取笔记详情
note_data = state.get('note', {}).get('noteDetailMap', {})
if note_data:
for key, note in note_data.items():
note_info = note.get('note', {})
result["title"] = note_info.get('title', '') or note_info.get('desc', '')[:50]
result["uploader"] = note_info.get('user', {}).get('nickname', '')
# 视频信息
video = note_info.get('video', {})
if video:
# 获取视频URL
media = video.get('media', {})
stream = media.get('stream', {})
# 尝试多种格式
for quality in ['h264', 'h265', 'av1']:
streams = stream.get(quality, [])
if streams:
for s in streams:
backup_urls = s.get('backupUrls', [])
if backup_urls:
result["video_url"] = backup_urls[0]
break
master_url = s.get('masterUrl')
if master_url:
result["video_url"] = master_url
break
if result["video_url"]:
break
# 如果还没找到,尝试consumer
if not result["video_url"]:
consumer = video.get('consumer', {})
origin_url = consumer.get('originVideoKey')
if origin_url:
result["video_url"] = f"https://sns-video-bd.xhscdn.com/{origin_url}"
# 封面
image_list = note_info.get('imageList', [])
if image_list:
result["cover_url"] = image_list[0].get('urlDefault')
if result["title"] or result["uploader"]:
result["success"] = True
break
except json.JSONDecodeError:
pass
# 方式2: 从 meta 标签提取
if not result["success"]:
title_match = re.search(r'<meta[^>]+property="og:title"[^>]+content="([^"]+)"', html)
if title_match:
result["title"] = title_match.group(1)
author_match = re.search(r'<meta[^>]+name="author"[^>]+content="([^"]+)"', html)
if author_match:
result["uploader"] = author_match.group(1)
video_match = re.search(r'<meta[^>]+property="og:video"[^>]+content="([^"]+)"', html)
if video_match:
result["video_url"] = video_match.group(1)
if result["title"]:
result["success"] = True
# 方式3: 从页面HTML中直接搜索视频URL
if not result["video_url"]:
video_patterns = [
r'"originVideoKey":"([^"]+)"',
r'"masterUrl":"([^"]+)"',
r'https://sns-video[^"]+\.mp4[^"]*',
]
for pattern in video_patterns:
match = re.search(pattern, html)
if match:
video_url = match.group(1) if '(' in pattern else match.group(0)
if not video_url.startswith('http'):
video_url = f"https://sns-video-bd.xhscdn.com/{video_url}"
result["video_url"] = video_url
break
return result
except Exception as e:
print(f" ⚠️ 获取笔记信息失败: {e}")
return result
def download_video(self, video_url: str, output_path: Path, progress_callback=None) -> bool:
"""下载视频到指定路径"""
headers = {
'User-Agent': self.pc_user_agent,
'Referer': 'https://www.xiaohongshu.com/',
'Accept': '*/*',
'Accept-Language': 'zh-CN,zh;q=0.9',
}
try:
response = requests.get(video_url, headers=headers, stream=True, timeout=60)
if response.status_code in (200, 206):
total_size = int(response.headers.get('content-length', 0))
downloaded = 0
output_path.parent.mkdir(parents=True, exist_ok=True)
with open(output_path, 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
downloaded += len(chunk)
if progress_callback and total_size > 0:
progress_callback(downloaded, total_size)
return True
else:
print(f" ⚠️ 下载失败,状态码: {response.status_code}")
return False
except Exception as e:
print(f" ⚠️ 下载视频时出错: {e}")
return False
def download(self, url: str, output_path: Path) -> bool:
"""
下载小红书视频的完整流程
Args:
url: 小红书笔记URL
output_path: 输出文件路径
Returns:
bool: 下载是否成功
"""
print(" 📕 检测到小红书链接,使用专用下载器...")
# Step 1: 获取笔记信息
note_info = self.get_note_info(url)
if not note_info.get("success"):
print(" ❌ 无法获取笔记信息")
return False
print(f" 📝 标题: {note_info.get('title', 'N/A')[:50]}...")
print(f" 👤 作者: {note_info.get('uploader', 'N/A')}")
video_url = note_info.get("video_url")
if not video_url:
print(" ❌ 这可能是图文笔记,不是视频笔记")
return False
# Step 2: 下载视频
print(f" 📥 开始下载视频...")
success = self.download_video(video_url, output_path)
if success:
file_size = output_path.stat().st_size / (1024 * 1024) # MB
print(f" ✅ 视频下载成功: {file_size:.2f} MB")
return success
def download_xiaohongshu_video(url: str, output_path: str) -> bool:
"""
下载小红书视频的便捷函数
Args:
url: 小红书笔记URL
output_path: 输出文件路径
Returns:
bool: 下载是否成功
"""
downloader = XiaohongshuDownloader()
return downloader.download(url, Path(output_path))
if __name__ == "__main__":
# 测试代码
if len(sys.argv) < 2:
print("Usage: python xiaohongshu_downloader.py <xiaohongshu_url> [output_path]")
sys.exit(1)
url = sys.argv[1]
output = sys.argv[2] if len(sys.argv) > 2 else "./xiaohongshu_video.mp4"
success = download_xiaohongshu_video(url, output)
sys.exit(0 if success else 1)
🎬 视频专家分析报告
📋 基本信息
| 项目 | 内容 |
|---|---|
| 视频 ID | {video_id} |
| 来源 URL | {url} |
| 分析时间 | {analysis_date} |
| 视频大小 | {video_size_mb} MB |
| 场景数量 | {scene_count} 个 |
| 转录语言 | {transcription_language} |
| 字幕段数 | {transcription_segments} 段 |
---
🎯 分析方法论
本报告基于 Walter Murch 剪辑六法则 和 动态权重评分系统:
核心哲学
情感 > 故事 > 节奏 > 视线追踪 > 2D平面 > 3D空间
>
一个情感真挚但画面略抖的镜头,优于一个画面完美但内容空洞的镜头。
五维评分体系 (1-10分制)
| 维度 | 权重 | 评估要点 |
|---|---|---|
| 美感 (Aesthetic) | 20% | 构图(三分法)、光影质感、色彩和谐度 |
| 可信度 (Credibility) | 20% | 表演自然度、物理逻辑、无出戏感 |
| 冲击力 (Impact) | 20% | 视觉显著性(Saliency)、动态张力 |
| 记忆度 (Memorability) | 20% | 独特视觉符号(Von Restorff效应)、金句、趣味性 |
| 趣味度 (Fun/Interest) | 20% | 参与感、娱乐价值、社交货币潜力 |
场景类型与权重适配
- TYPE-A Hook/高能型: IMPACT 40% + MEMORABILITY 30% + SYNC 20%
- TYPE-B 叙事/情感型: CREDIBILITY 40% + MEMORABILITY 30% + AESTHETICS 20%
- TYPE-C 氛围/空镜型: AESTHETICS 50% + SYNC 30% + IMPACT 20%
- TYPE-D 商业/展示型: CREDIBILITY 40% + MEMORABILITY 40% + AESTHETICS 20%
筛选决策规则
| 等级 | 标准 | 用途 |
|---|---|---|
| 🌟 MUST KEEP | 加权总分 > 8.5 或 任意单项 = 10 | 核心素材,极致长板 |
| 📁 USABLE | 7.0 <= 加权总分 < 8.5 | 过渡素材,辅助叙事 |
| 🗑️ DISCARD | 加权总分 < 7.0 或存在致命瑕疵 | 建议舍弃 |
---
📝 视频内容概述
转录文本
{transcript_text}内容主题分析
(请在完成场景评分后,由分析师填写)
- 视频类型: TODO (广告/剧情/Vlog/教程/其他)
- 核心主题: TODO
- 目标受众: TODO
- 情感基调: TODO (励志/搞笑/温馨/悬疑/其他)
- 商业意图: TODO (如有)
---
🎞️ 场景详细分析
场景列表概览
| 场景编号 | 类型 | 加权得分 | 筛选建议 | 片段路径 |
|---|
{scene_list_table}
各场景详细评估
(以下部分需要基于帧图片进行专业分析后填写)
{detailed_scene_evaluations}
---
⭐ 精选片段推荐
入选精选文件夹的片段
(完成评分后,得分 >= {best_threshold} 的片段将自动入选)
入选标准: 加权总分 >= {best_threshold} 或 单项得分 = 10 (极致长板)
精选片段位置: scenes/best_shots/
{best_shots_table}
最佳 Hook 候选
TODO: 基于 IMPACT 和 MEMORABILITY 得分,推荐最适合作为开场 Hook 的场景
最佳情感片段
TODO: 基于 CREDIBILITY 得分,推荐最能建立共情的场景
最佳视觉片段
TODO: 基于 AESTHETICS 得分,推荐视觉最精美的场景
---
📊 整体影片评价
综合评分
| 评价维度 | 得分 (1-10) | 评价说明 |
|---|---|---|
| 整体美感 | TODO | TODO |
| 叙事连贯性 | TODO | TODO |
| 情感共鸣度 | TODO | TODO |
| 商业转化力 | TODO | TODO (如适用) |
| 病毒传播潜力 | TODO | TODO |
| 综合得分 | TODO | - |
优势分析
TODO: 总结视频的核心优势和亮点
改进建议
TODO: 基于 Walter Murch 法则提出具体的改进建议
最终 verdict
| 评价项 | 结论 |
|---|---|
| 是否值得保留 | TODO |
| 推荐使用场景 | TODO |
| 二次创作建议 | TODO |
---
📁 文件结构
{video_output_dir_name}/
├── {video_id}.mp4 # 完整视频
├── {video_id}.m4a # 音频文件
├── {video_id}.srt # 字幕文件
├── {video_id}_transcript.txt # 转录文本
├── scene_scores.json # 评分模板
├── {video_id}_analysis_report.md # 基础报告
├── {video_id}_detailed_analysis.md # 本详细报告
├── scenes/ # 场景片段
│ ├── {video_id}-Scene-001.mp4
│ ├── {video_id}-Scene-002.mp4
│ ├── ...
│ └── best_shots/ # ⭐ 精选片段
│ ├── (评分后自动复制入选片段)
├── frames/ # 场景预览帧
│ ├── {video_id}-Scene-001.jpg
│ └── ...---
📝 后续步骤
1. 查看帧图片: 打开 frames/ 目录查看每个场景的预览 2. 完成评分: 编辑 scene_scores.json,为每个场景填写评分和评价 3. 计算排名: 运行 python3 scripts/scoring_helper_enhanced.py scene_scores.json calculate 4. 生成精选: 运行 python3 scripts/scoring_helper_enhanced.py scene_scores.json best 5. 完善报告: 根据评分结果,完善本报告的 TODO 部分
---
本报告由 Video Expert Analyzer 自动生成 基于 Walter Murch 剪辑六法则 和 动态权重评分系统
Related skills
FAQ
What model does video-expert-analyzer require?
A multimodal vision-capable model such as Gemini 3.0, Kimi 2.5, or Claude in Agent mode; text-only models cannot score frames.
How are scenes selected?
Each scene is scored on five dimensions; weighted scores of 8.5+ are MUST KEEP, 7.0 to 8.5 are USABLE, and below 7.0 are DISCARD.