
Gpt Image2 Ppt
- 114 installs
- 1.1k repo stars
- Updated August 2, 2026
- juneyaooo/gpt-image2-ppt-skills
Helps with ai & agent building tasks.
About
gpt-image2-ppt is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted coding.
- gpt-image2-ppt
- AI & Agent Building
- AI-coding skill
Gpt Image2 Ppt by the numbers
- 114 all-time installs (skills.sh)
- +7 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #3,964 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/juneyaooo/gpt-image2-ppt-skills --skill gpt-image2-pptAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 114 |
|---|---|
| repo stars | ★ 1.1k |
| Last updated | August 2, 2026 |
| Repository | juneyaooo/gpt-image2-ppt-skills ↗ |
What it does
Helps with ai & agent building tasks.
Files
gpt-image2-ppt -- 用 gpt-image-2 生成 PPT
把一份 markdown 大纲(或 slides_plan.json)+ 一种视觉风格,直接喂给 OpenAI 官方 Images API(gpt-image-2),逐页出图,最后打包成 16:9 .pptx。
可用风格
| 风格 ID | 一句话定位 | 适用场景 |
|---|---|---|
gradient-glass | Apple Vision OS / Spatial Glass | AI 产品发布、技术分享、创意提案 |
clean-tech-blue | Stripe / Linear 级蓝白 | 融资路演、商业计划书、企业战略 |
vector-illustration | 复古矢量插画 + 黑描边 | 教育培训、品牌故事、社区分享 |
editorial-mono | Kinfolk / Monocle 编辑设计 | 品牌发布、文化访谈、读书分享 |
dark-aurora | Linear / Vercel 深色霓虹 | AI 产品、开发者工具、技术分享 |
risograph | Riso 双套色印刷 + 网点纹理 | 创意工作室、文创品牌、独立 zine |
japanese-wabi | 无印 / 原研哉式侘寂 | 茶道、生活方式、奢侈品、文化讲座 |
swiss-grid | Bauhaus / Vignelli 国际主义网格 | 学术报告、博物馆展陈、严肃汇报 |
hand-sketch | Sketchnote / 白板手绘 | 工作坊、产品 brainstorming、培训 |
y2k-chrome | Y2K 千禧液态金属 + 蝴蝶贴纸 | 潮牌、文娱、品牌联名、Z 世代营销 |
abstract-art-showcase | 黑白极简、艺术展览感、超大字体和抽象画面并置 | 艺术策展、作品集、品牌调性展示 |
coal-industry-business-company-profile | 工业棕黑、粗重标题、结构线和硬朗图标 | 能源、制造业、重资产公司介绍 |
college-candy-aesthetics-infographics | 糖果色、校园感、圆润信息图和轻快装饰 | 教育、校园活动、轻量数据科普 |
creative-agency | 创意机构气质、强视觉拼贴、鲜明版式节奏 | Agency 提案、品牌方案、创意汇报 |
culinary-innovation | 餐饮创新感、食材摄影、暖色块和杂志式排版 | 餐饮品牌、食品创新、菜单/新品发布 |
data-science-consulting | 数据咨询蓝灰、模块化布局、图表和技术感信息层级 | 数据分析、AI 咨询、企业数字化 |
mindfulness-in-the-classroom-breathing-techniques | 柔和心理健康配色、留白、圆角块和安静插画感 | 心理健康、课堂活动、呼吸训练课程 |
mind-maps-workshop-professional | 专业工作坊风、思维导图节点、清晰流程结构 | 培训工作坊、方法论、团队共创 |
meeting-agenda | 会议议程感、干净网格、强信息分组和商务标题 | 例会、项目同步、管理层汇报 |
investment-company-business-plan | 投资机构质感、深浅对比、稳重商务版式 | 投资计划、基金介绍、商业计划书 |
indigenous-cultures | 文化纹样、自然色、手工质感和叙事型构图 | 文化课程、历史主题、公益教育 |
health-disparities-and-social-determinants-of-health-doctor-of-philosophy-phd-in-health-behavior-and-health-education | 公共健康学术风、理性网格、柔和医疗色和论文感层级 | 医学论文答辩、公共健康报告、教育研究 |
geometric-duotone-thesis | 双色几何、论文答辩感、斜切图形和强标题 | 学术答辩、研究报告、章节型内容 |
geometric-clinical-case | 几何医疗风、冷静配色、病例卡片和清晰分栏 | 临床病例、医疗培训、诊疗汇报 |
geometric-business | 商务几何块、稳健蓝绿调、简洁图表语言 | 商业计划、团队汇报、产品策略 |
formal-lavender-portfolio | 淡紫正式感、作品集留白、优雅细线和柔和版式 | 个人作品集、设计简历、专业展示 |
flowery | 花卉装饰、柔和色块、浪漫但有秩序的排版 | 生活方式、女性品牌、活动介绍 |
first-impressions | 第一印象主题、强封面视觉、人物/标题的戏剧化关系 | 面试培训、个人品牌、沟通课程 |
final-year-project-thesis-defense | 毕业设计答辩、学院派网格、清晰章节与数据页 | 毕业答辩、项目结题、研究展示 |
fashion-business-consulting-toolkit-aesthetic | 时尚咨询感、高级拼贴、杂志排版和中性色 | 时尚商业、品牌咨询、趋势报告 |
economic-impact-of-coronavirus | 经济影响报告风、严肃信息图、冷静色彩和数据叙事 | 宏观经济、政策分析、风险报告 |
eco-green-business-plan | 鼠尾草绿、自然材质摄影、环保商务与极简分屏 | 可持续商业、环保品牌、健康生活方式 |
所有可用风格都统一放在 styles/ 下,使用方式完全相同。需要查看风格封面展示时,读取 docs/distilled-styles.md。
风格选择原则:先根据内容场景在styles/里选择最贴近的一套。技术类可优先看dark-aurora/gradient-glass/data-science-consulting,商务类可优先看clean-tech-blue/editorial-mono/eco-green-business-plan/investment-company-business-plan,文化生活类可优先看japanese-wabi/vector-illustration/culinary-innovation/flowery,学术类可优先看swiss-grid/geometric-duotone-thesis/final-year-project-thesis-defense,工作坊与培训类可优先看hand-sketch/mind-maps-workshop-professional/mindfulness-in-the-classroom-breathing-techniques。
模板克隆模式
直接给 skill 一个 .pptx 模板,后续所有页都仿这个模板。
# 一行:自动渲染 + 模板分析 + 出图。需本机有可用 PPTX 渲染后端
python3 scripts/generate_ppt.py \
--plan slides_plan.json \
--template-pptx ./company-template.pptx \
--template-strict--template-strict 表示每页都把模板对应页作为 image reference 喂给 gpt-image-2,仿真度最高。
模板渲染:本机不需要操作 PowerPoint
skill 自带 render_template.py,把 .pptx 自动渲染成每页 PNG,存到 <cwd>/template_renders/<stem>/page-NN.png。
Agent 前置检查(模板克隆时必须做)
在跑任何 --template-pptx 命令之前,你必须先检查本机是否有可用 PPTX 渲染后端。
检查方式:
- 首选:在 skill 目录运行
python3 scripts/render_template.py --check。它会验证后端是否真的可执行,而不是只看路径是否存在。 - macOS:优先检查
/Applications/Keynote.app且 AppleScript 可执行;否则检查libreoffice --version || soffice --version - Windows:优先检查本机 PowerPoint COM 可启动;否则检查
libreoffice --version/soffice --version - Linux / 兼容层:检查
libreoffice --version || soffice --version,不要只用which
注意:鸿蒙 / Termux / 容器 / 特殊架构环境可能看起来像 Linux,但不能假设 Linux aarch64 的 LibreOffice 二进制可运行;必须以 render_template.py --check 或 soffice --version 的实际执行结果为准。不要把 aspose-slides 当默认兜底,它在很多移动/特殊 Python 环境没有可安装 wheel。
如果都没有可用后端,先告知用户模板渲染需要安装可执行的 LibreOffice,或让用户在桌面端手动把模板每页导出为 page-01.png、page-02.png 后通过 --template-images 传入。可选安装命令:
| 平台 | 安装命令 |
|---|---|
| Windows | winget install LibreOffice.LibreOffice |
| macOS | brew install --cask libreoffice |
| Linux (Debian/Ubuntu) | sudo apt-get install -y libreoffice |
| Linux (Fedora/RHEL) | sudo dnf install -y libreoffice |
| Linux (Arch) | sudo pacman -S --noconfirm libreoffice-fresh |
装完再次检查,确认存在可用渲染后端再继续后续流程。
注意:Windows 上winget是 Win10/11 自带,会弹 UAC 确认框,需要用户点确认;macOS 上brew需要先安装 Homebrew。
render_template.py 的渲染后端按优先级自动挑: 1. Windows:PowerPoint COM(本机有 Office 时优先,直出 PNG,跳过 PDF 步骤)> LibreOffice 2. macOS:Keynote AppleScript(本机有 Keynote 时优先,直出 PNG)> LibreOffice 3. Linux / 兼容层:通过 --version 探测确认可运行的 LibreOffice / soffice 命令 4. PDF -> PNG 走 pymupdf(已在 requirements);没装就用 pdf2image + poppler
跑 generate_ppt.py --template-pptx ... 时如果省略 --template-images 会自动调一次渲染;也可以手动先跑一次:
python3 scripts/render_template.py company-template.pptx
# -> <cwd>/template_renders/company_template/page-01.png ... page-NN.png仿模板的两层缓存
| 资料 | 路径 | 用途 |
|---|---|---|
| 模板每页 PNG | <cwd>/template_renders/<stem>/page-NN.png | 本机渲染后端一次渲染长期复用 |
| 模板风格分析 | <cwd>/template_cache/<sha256>.json 或手写 template_profile.json | 多模态 agent 自己看图生成;纯文本 agent 才需要外挂 vision |
| 生成产物 | <cwd>/outputs/<timestamp>/ | 每次新跑都新目录 |
三者都在调用者 cwd 下,与项目自然同进退;建议把 template_renders/、template_cache/、outputs/ 加进项目的 .gitignore。
*模板看图分析(让 agent 自己判断要不要配 `VISION_`)**:
- 当前 code agent 本身是多模态模型(例如 Claude Code 的多模态 Claude、Codex 的多模态 GPT):不需要额外配置
VISION_*。agent 直接读取template_renders/<stem>/page-*.png,按template_analyzer.py的TemplateProfile结构生成template_profile.json,再用--template-profile template_profile.json传给generate_ppt.py。如果要配合--template-strict,每个 layout 里要写reference_image(模板 PNG 的绝对路径或可访问路径)。 - 当前 code agent 是纯文本模型(例如只接入 DeepSeek 文本模型):它看不了模板截图,需要额外配置
VISION_BASE_URL/VISION_API_KEY/VISION_MODEL_NAME,让template_analyzer.py调一个独立的 OpenAI 兼容多模态端点做模板分析。
vision 分析与图片生成的 gpt-image-2 永远解耦——换 vision provider 不影响出图路径。
安装
git clone git@github.com:JuneYaooo/gpt-image2-ppt-skills.git
cd gpt-image2-ppt-skills
bash install_as_skill.sh --target claude # Claude Code
# 或
bash install_as_skill.sh --target codex # Codex
# API 直连所需密钥优先通过 agent 配置 / 系统环境变量注入环境变量注入(API 直连时)
不要把本 skill 的密钥写进调用者业务项目根目录的 .env,也不要为了出图去读取用户项目里的通用 .env。环境变量建议按 agent 框架的标准方式注入:
- 通用 / CI / 服务器:用系统环境变量、Docker Compose
environment/env_file、Kubernetes Secret、CI Secret 等注入。 - Claude Code:用用户级
~/.claude/settings.json或项目级.claude/settings.local.json注入环境变量;命令行环境变量优先级最高。 - OpenClaw / 自定义 Agent:用框架配置里的
apiKey/ env reference 引用系统环境变量,避免把 key 明文写进项目配置。 - 本地 standalone CLI fallback:可以设置
GPT_IMAGE2_PPT_ENV=/path/to/private.env,或使用 skill 安装目录下的.env;这只是备用方式,不是业务项目.env。
API 直连需要这些变量:
OPENAI_BASE_URL=https://api.openai.com # 或任意 OpenAI 兼容中转站
OPENAI_API_KEY=sk-...
GPT_IMAGE_MODEL_NAME=gpt-image-2
GPT_IMAGE_QUALITY=high # low / medium / high / auto
# 可选:模板克隆模式的 vision 分析 backend。
# 多模态 agent / 原生 Codex 可自己看图生成 --template-profile,不需要下面这组。
# 只有纯文本 agent(如 DeepSeek 文本模型)才需要外挂下面这组。
# 不内置默认 endpoint,请填你自己信任的服务,否则就别填。
# VISION_BASE_URL=https://your-openai-compatible-relay.example.com/v1
# VISION_API_KEY=sk-...
# VISION_MODEL_NAME=gemini-3.1-pro-preview # 或 gpt-4o / claude-3.5-sonnet 等任意多模态 SKU安全提示:脚本只读取当前进程环境、平台注入的gpt-image2-ppt_*变量、显式GPT_IMAGE2_PPT_ENV,以及 skill 安装目录下的.envfallback。脚本不会向上递归读取调用者项目目录里的.env,避免误吃业务项目密钥。
如果你就是 Codex agent(原生 image_generation 出图 — 推荐)
如果你自己就是 Codex(正在运行本 skill 的 agent 就是 Codex CLI / Codex TUI),并且当前环境提供 image_generation tool 和 ChatGPT 登录态,此时不要用 `generate_ppt.py` 或 `--backend codex` 负责出图,直接用原生工具生成图片,最后只复用本仓库的 md 转换 / PPTX 打包逻辑即可。
关键边界:Python 脚本运行在子进程里,拿不到当前 agent 会话里的原生 tool。generate_ppt.py --backend codex 能做的只有再启动一个 codex exec 子进程,让另一个 Codex 去出图;它不是“复用当前 Codex 的 image_generation tool”。所以当前 agent 已经能原生出图时,出图动作必须由 agent 本身完成,而不是交给 generate_ppt.py。
如何判断
你能访问 image_generation tool,并且不需要手动配 OPENAI_API_KEY 就能出图——满足这两个条件就走原生路径。若当前 Codex 会话没有这个 tool,就按普通 agent 处理:走 API 直连、--backend codex 备用后端,或让用户补齐环境。
出图流程(Codex 原生路径)
1. 准备 slides 数据
如果还没有 slides_plan.json,先按下面「生成流程」第 2-3 步写 slides_plan.md → python3 scripts/md_to_plan.py ... 转 json。
2. 读风格模板
读 styles/<id>.md,取 ## 基础提示词模板 section 作为 base prompt。
3. 构造每页 prompt
参考 generate_ppt.py 的 generate_prompt() 逻辑,核心规则:
- 封面(cover/slide 1):标题/副标题为视觉焦点
- 数据页(data/最后一页):突出关键数字、对比或结论
- 内容页(content/其余页):按层级、对齐、留白结构化呈现
- 所有文字必须简体中文,字体用思源黑体/苹方,严禁草书/艺术字
- 16:9 横版宽屏(landscape, widescreen),prompt 里明确说"宽度明显大于高度、绝对不要方图"
{style 基础提示词模板}
---
现在请生成本组中的【{封面页/内容页/数据页}】,{对应 hint}
本页要呈现的内容如下(请按本风格美学重新设计版式):
{slide content}
【强制语言与字体要求】
1. 所有文字必须使用简体中文,严禁英文(专有名词除外)
2. 中文字体使用思源黑体或苹方,严禁草书、艺术字
3. 标题粗体,正文常规,字号对比清晰
【画面比例 — 强制】16:9 横版宽屏 (landscape, widescreen),宽度明显大于高度,绝对不要方图或竖图。4. 调 image_generation tool 出图
对每页调你的 image_generation tool:
prompt: 上面拼好的完整 promptoutput_format:png- 将返回的图片保存到
outputs/<timestamp>/images/slide-NN.png(NN 为两位页码)
可以并发(建议 ≤4 并发,避免限流)。
5. 打包 PPTX
如果本 deck 没有外部真实图片对象,可以用下面的简易整页 PNG 打包。如果任一页有 `external_image` / `image_overlay` / `external_image_placeholder` 且指向真实图片,不能用这个简易打包片段,否则真实图片会被合进整页背景 PNG,用户无法在 PowerPoint 里单独选中拖动。此时必须走 generate_ppt.py 的标准打包逻辑,或在已有 session 中调用 generate_pptx(..., metadata=metadata),让真实图片作为独立 picture object 叠在背景上。
python3 -c "
from pptx import Presentation
from pptx.util import Inches
prs = Presentation()
prs.slide_width = Inches(13.333)
prs.slide_height = Inches(7.5)
blank = prs.slide_layouts[6]
import os, glob
for p in sorted(glob.glob('outputs/<timestamp>/images/slide-*.png')):
slide = prs.slides.add_slide(blank)
slide.shapes.add_picture(p, 0, 0, width=prs.slide_width, height=prs.slide_height)
prs.save('outputs/<timestamp>/<title>.pptx')
print('done')
"外部真实图片页的正确 PPTX 结构应是:
- 第 1 层:
images/slide-XX.png作为整页背景图(包含模型生成的背景和文字)。 - 第 2 层:
source指向的真实图片作为独立 PPT picture object,按slide_spec坐标贴在背景上,可在 PowerPoint 里选中、拖动、缩放。
模板克隆模式(Codex 原生路径)
你自己就是多模态 agent——直接 Read 模板每页 PNG 抽取视觉风格,写成 template_profile.json(schema 见 template_analyzer.py 里的 TemplateProfile,每个 layout 写上 reference_image),然后按上面流程出图时把对应模板页作为 reference image 传给 image_generation tool。
*不需要配 `VISION_`**——你就是 vision。
与下面「--backend codex」的区别
| 原生路径(本节) | --backend codex | |
|---|---|---|
| 适用场景 | 你就是 Codex agent | 你是 Claude Code / 其他 agent,借用本机 codex CLI |
| 调用方式 | 直接调 image_generation tool | spawn codex exec --full-auto 子进程 |
| 出图层数 | 1 层 | 2 层(agent → python → codex exec) |
| 速度 | 几秒/张 | 30-60s/张 |
| 可靠性 | tool 参数精确 | 自然语言 relay,偶发失败 |
| 需要 API Key | 不需要 | 不需要 |
---
可选:走 codex CLI 出图(--backend codex,非 Codex caller 用)
如果你就是 Codex agent,不要走这条路——用上一节的「原生路径」代替。
当你用 Claude Code / OpenClaw / 其他 agent 运行本 skill,但本机装了 codex CLI 且已登录(codex login),可以借用它的凭据出图,省掉配 OPENAI_API_KEY:
python3 scripts/generate_ppt.py --plan slides_plan.json --style styles/editorial-mono.md --backend codex默认后端仍是 openai(直调 API,快、并发稳、每页 3-10s)。--backend codex 是逃生口,适合"只跑 1-2 张图试水、不想配 key"的场景。
Tradeoffs:
- ✅ 不需要在本 skill 配
OPENAI_API_KEY - ⚠️ 慢:每页多一层 agent loop,单页 30-60s+,10 页可能 5-10 分钟
- ⚠️ 计费不变:gpt-image-2 是按图计费,不在 ChatGPT 订阅内,codex 只是代你刷额度
- ⚠️ 可控性差:aspect_ratio / quality / reference_image 靠自然语言指令让 codex 转发,偶发失败
相关 env(都可选):
CODEX_CMD="codex exec --full-auto" # 覆盖 codex 调用方式(默认这串)
CODEX_IMAGE_MODEL=gpt-image-2 # 传给 codex 的目标模型
CODEX_TIMEOUT_SECS=900 # 单页超时
GPT_IMAGE_BACKEND=codex # 不想每次敲 --backend 就设这个模板克隆的 vision 分析同理——当 caller agent 自己是多模态时(Claude Code / 多模态 codex),可以直接 Read 模板 PNG 抽取风格,不用配 VISION_*;只有 caller agent 是纯文本模型时才需要外挂 vision provider。
生成流程(指定风格)
先 md 后 json:md 给人看、方便 diff / review / 改文案;json 由 md 派生,喂给 generate_ppt.py,标为 generated,不手改。
1. 用户给一份大纲 / 已有的 slides_plan.json 2. Agent 按下面 md 规范写一份 slides_plan.md,与用户确认文案: ````markdown --- title: MediWise Health Suite 商业计划书 ---
1. [cover] MediWise Health Suite
副标题:家庭健康管理智能平台 年份:2026
2. [content] 市场痛点:健康管理的两类割裂
痛点一:高频无深度 ...
6. [data] 效率对比:使用 MediWise 前后
... ````
- h2 格式:
## N. [page_type, layout=layout-05] 本页标题行 N.可省(按出现顺序自动编号);[page_type]可省(默认content);layout=只在模板克隆模式需要page_type:cover/content/data- h2 标题行 → json 里
content的第一行;下面的正文 → 正文
3. 用户 OK 后,转 json:
python3 scripts/md_to_plan.py slides_plan.md -o slides_plan.json4. 选风格:从 styles/ 里挑一个,对应 styles/<id>.md;需要视觉预览时先看 docs/distilled-styles.md 5. 构造 slide_spec(Agent 步骤):读 styles/<id>.md 的视觉规范,为 slides_plan.json 每页构造 slide_spec(每个元素的 type、content、position、style),写入每页的 slide_spec 字段。格式见下方"指哪改哪"章节 6. 调脚本:
python3 scripts/generate_ppt.py --plan slides_plan.json --style styles/editorial-mono.md7. 产物在 <cwd>/outputs/<timestamp>/:
images/slide-XX.png-- 每页 PNG(16:9,1536x864)prompts.json-- 每页用到的完整 prompt(便于复盘 / 二次微调)metadata.json-- slide_spec 版本历史(支持精确编辑和回滚)<title>.pptx-- 16:9 PPTX;默认背景与文字是整页图片,通过external_image放入的真实图片会作为独立 PPT 图片对象叠加
外部真实图片贴入(推荐精确流程)
默认规则:用户提供真实图时,优先按原图保真后贴,不要让 gpt-image-2 重画这张图。使用 slide_spec 声明外部图片槽位:
这套流程只在元素声明了 type: "external_image" / image_overlay / external_image_placeholder 且 source 能解析到真实本地图片文件时启用。没有真实图片 source 的普通生成、模板克隆、纯占位布局和老的自由风格生成不受影响。
如果用户明确说“更重视画面融合效果,不需要一定贴原图 / 可以重绘 / 可以图生图”,可以走参考图模式:把图片作为 generation reference 输入给模型,而不是最终独立后贴。此时版面通常更融合,但不保证像素级保真,PPT 里也不会有可单独选中的原图对象。
参考图模式可用 type: "image_reference",或在 external_image 上显式写 render_mode: "reference" / preserve_original: false:
{
"elements": {
"mood_reference": {
"type": "image_reference",
"source": "/absolute/path/to/photo.png",
"purpose": "只作为视觉参考,允许模型融合重绘,不作为独立 PPT 图片对象后贴"
}
}
}不要对高精度素材使用参考图重绘:医疗影像、病理图、诊断依据、实验/工程读数、财务表格、论文图表、法律证据截图、产品 UI 精确截图等都应默认走 external_image 保真后贴,并在交付前提示用户核对。
{
"elements": {
"hero_photo": {
"type": "external_image",
"source": "/absolute/path/to/photo.png",
"layout_intent": "auto",
"tailor_to_asset": true,
"slot_strategy": "fit-within",
"fit": "contain",
"slot": {
"padding": 0.012,
"bleed": 0,
"fill": "#F7F7F5",
"mask_placeholder": false,
"sanitize_background": false,
"draw_frame": false,
"outline_width": 0,
"skeleton_canvas_fill": "transparent",
"skeleton_fill": "transparent",
"skeleton_shape": "corners",
"skeleton_outline": "#000000",
"skeleton_outline_width": 2,
"skeleton_ticks": false
}
}
}
}生成时脚本会自动做四件事:
1. 先规划槽位,再生成 skeleton。脚本会优先读取模板 profile 里的 external_image_slots,或根据模板摘要推断“左图右文 / 右图左文 / 底部图表 / 中央主视觉”等候选区域;然后结合本页文字量、标题长度、真实图片数量、真实图片宽高比和素材分类(照片 / 图表 / 文档 / 架构图等)打分选位,先产出 position / computed_bbox / auto_layout_reason / layout_planning_profile。skeleton 只是把这个规划结果画给模型看,不负责临时想位置。 2. 读取 source 真实图片尺寸。如果声明了 tailor_to_asset: true 或 slot_strategy: "fit-within",脚本会把 position 当作“可用区域”,按真实图片宽高比在其中计算 computed_bbox;这个 bbox 会同时用于 prompt、骨架参考图和最终 PPTX 贴图。这样不是生成后再临时缩放,而是在 gpt-image-2 出图前就量体裁衣。 3. 在 outputs/<timestamp>/references/slide-XX-asset-skeleton.png 生成一张透明画布的角标骨架参考图,把最终真实图会覆盖的 final_image_rect_px 标出来,并作为 reference image 传给 gpt-image-2。这只用于引导模型不要把关键信息放进该区域;每页只调用一次 `gpt-image-2`,不是先生成一张再用骨架二次重生。 4. 打包 PPTX 时,按同一个 bbox 用 python-pptx 把 source 指向的真实图片作为独立图片对象贴入。默认不画额外框,也不铺遮罩;也不会默认清理 gpt-image-2 生成图。真实图片不应被合成进 images/slide-XX.png,否则用户无法在 PPT 里单独拖动。
如果生成后发现真实图片槽位压到大面积文字或图形,优先按下面两种方式处理:
1. 预生成前微调槽位:改 slide_spec.elements.<id>.position(或 anchor / padding),让真实素材移动到更干净的区域,再重新生成该页。只要槽位位置变了,就必须重新生成背景;不要只在 PPTX 里移动最终图片,否则背景内容仍可能压住新位置。 2. 二次版式修复:当第一页整体风格已经满意,只是局部内容压入槽位时,用 --edit SLIDE --external-slot-repair。该模式会同时把当前页和新的 asset-skeleton 作为 reference image 传入,要求模型保持原风格但重排文字/图形,让槽位空出来。
示例:把真实图区域右移并重排当前页:
python3 scripts/generate_ppt.py \
--session outputs/20260529_120000 \
--edit 3 \
--external-slot-repair \
--element-updates '{"hero_photo":{"position":[0.60,0.16,0.32,0.60],"anchor":"center"}}'同时,每个包含外部真实图片的页面都会在 outputs/<timestamp>/external_image_trace/slide-XX/ 记录完整中间链路:
step1-real-on-blank.png:把真实图片按最终坐标和大小先贴到空白页上,用来确认“真实图最终应该出现在这里”。step2-reference-outline-blank.png:去掉真实图片,只保留给gpt-image-2的角标定位参考图;默认透明画布,不铺白底。定位区域使用final_image_rect_px,也就是最终真实图实际会覆盖的像素矩形,而不是外层 slot。这张图同时会复制到references/slide-XX-asset-skeleton.png并作为 image reference 输入。step3-generated-background.png:gpt-image-2根据 step2 reference 生成的背景图,尚未贴入真实图片。step3-image2-raw.png/step3-sanitized-background.png:只有显式设置sanitize_background: true时才会出现,用于兜底清理模型已经画出的占位块;默认不启用。step4-final-overlay-preview.png:在 step3 上按 step1 同一坐标贴回真实图片的预览图,用来和 PPTX 最终效果对照。manifest.json:记录每个外部图片的source、slot_bbox_norm、inner_rect_norm、reference_rect_px、final_image_rect_px、padding、bleed、素材尺寸和比例。
关键原则:
- reference skeleton 负责“告诉模型哪里不要放关键信息”,不是要求模型画白色空框。
- 模板克隆 /
--template-strict与外部真实图片槽位同时存在时,生成阶段会同时传入两张 reference:模板页用于学习风格,asset-skeleton用于标记后贴图片覆盖区;不要二选一,否则模型只看模板页时不会知道最终图片位置。 - 外部真实图片页不能简单堆叠“模板照片/图片区 prompt”和“留空 prompt”。构造 prompt 前必须先消解冲突:模板只提供配色、字体、网格、节奏和装饰语言;模板里的照片区、图片框、全幅图、裁切图等都视为已被真实图片槽位替代。
- 构造
slide_spec时,position不能随便选,也不要默认固定右侧。模板克隆模式下,优先让模板 profile 提供external_image_slots,或从模板布局摘要推断候选图片区,再根据本页文字量、真实图片数量、真实图片比例和页面类型打分:文字多时优先保证标题/正文阅读区;图片多或比例极端时优先用模板原有分栏/上下布局;无法共存时减少装饰和假视觉主体,而不是让真实图压内容。 - 推荐先写
layout_intent: "auto",不手写position。脚本会读取模板候选槽位、真实图尺寸、文字量、图片数量和比例来选择位置,写入position/computed_bbox/auto_layout_reason;只有自动规划不理想时才用人工position覆盖。 - 使用
layout_intent: "auto"时,slide_spec.layout尽量写清图片区方向和结构,例如"左文右图 / right image / right landscape photo / 底部图片 / portrait rail"。自动规划会优先从这些方向性描述和模板候选槽位推断真实图区域;描述过泛时,重文本页可能退化到保守的小图槽位。 position默认可作为“允许使用的区域”;computed_bbox才是脚本算出的最终真实图片槽位。- 如果你想完全固定坐标,不要量体裁衣,直接写
bbox,或设置slot_strategy: "exact"/tailor_to_asset: false。 - 最终精度由
python-pptx坐标保证,不由模型画框保证。 skeleton_outline_width只控制给模型看的骨架参考图;outline_width/draw_frame才控制最终 PPT 里是否出现可见框。- prompt 中不写归一化坐标;位置通过透明角标 reference 图控制,避免模型错误解释数值坐标,也避免闭合框被理解成真实图片框。
skeleton_canvas_fill/skeleton_fill/skeleton_outline/skeleton_outline_width/skeleton_shape只影响参考骨架图。默认skeleton_canvas_fill: "transparent"且skeleton_shape: "corners",只在final_image_rect_px四角画短角标,不画闭合矩形、不铺白底,避免模型把 reference 理解成真实图片框或白色占位块。需要旧行为时显式设skeleton_shape: "outline"。sanitize_background/mask_placeholder不是推荐路径;它们只适合清掉极轻微的占位痕迹,不适合解决文字或图形大面积压住槽位的问题。遇到压内容,按“微调槽位并重生”或“二次版式修复”处理。- 默认最终不画外框、不插入任何底色矩形,只贴真实图片;
mask_placeholder会被忽略,避免在真实图下方产生可拖动白底。只有显式设置draw_frame: true且outline_width > 0时,才会额外画无填充边框。 - 推荐默认:
tailor_to_asset: true+slot_strategy: "fit-within"+fit: "contain"+ 少量padding。这会完整保留原图,并让骨架图从一开始就是按素材比例预留的。 fit: "contain"保留完整图片;fit: "cover"会居中裁剪图片来铺满槽位。- 如果要“严丝合缝”无白边,使用
fit: "cover"、padding: 0、outline_width: 0,并设置很小的bleed(例如0.0015-0.003)让图片比槽位多铺出约 1-3px,抵消 PowerPoint / Quick Look 渲染取整和抗锯齿缝隙。 - 如果想“预留一点但不要莫名空隙”,不要靠后期
contain硬塞进一个比例不匹配的大框;应使用fit-within先算好比例匹配的槽位,再用很小的padding做设计留白。 - 交付前必须抽查
external_image_trace/slide-XX/manifest.json:重点看final_image_rect_px是否符合视觉预期,auto_layout_reason是否选中了正确的左右/上下区域。若横图被算成小缩略图、竖图过窄或位置压内容,先补清楚slide_spec.layout的方向意图,或显式设置position,再重新生成该页背景并重新打包 PPTX。
整体原理图保存于 docs/external_image_overlay_logic.txt;修改外部真实图片链路时,先对照这张图确认顺序仍是“模板/页面/素材画像 -> 候选槽位打分 -> 透明 skeleton -> 背景生成 -> 后贴真实图 -> 质检/有限修复”。
生成流程(模板克隆)
1. 拿到模板 .pptx(用户提供 / 内部模板库 / 网络下载) 2. (可选)先单独渲染并人工挑选----大模板(>15 页)建议先 python3 scripts/render_template.py xxx.pptx,再从 template_renders/<stem>/ 里挑 8-12 张代表页复制到 template_renders/<stem>_curated/,供 vision 分析。页数越精,layout 命中越准 3. 生成 slides_plan.md → 转 slides_plan.json(见指定风格流程第 2-3 步)。每页 slide_number / page_type (cover / content / data / 等) / content;想精准对位时在 h2 里加 layout=layout-NN(NN = 模板第 N 页 / 你期望对应的模板页编号) 4. 出图冒烟。API 直连 / 非 Codex 原生路径跑 generate_ppt.py:
python3 scripts/generate_ppt.py \
--plan slides_plan.json \
--template-pptx xxx.pptx \
--template-images template_renders/xxx_curated \
--template-strict --slides 1如果当前 agent 自己是多模态模型,也可以先看模板 PNG 写出 template_profile.json,再这样跑,完全不需要 VISION_*:
python3 scripts/generate_ppt.py \
--plan slides_plan.json \
--template-profile template_profile.json \
--template-strict --slides 1先 --slides 1 出封面冒烟,效果 OK 再跑全量。
如果当前 agent 就是带原生出图能力的 Codex,不要用上面的 generate_ppt.py 命令负责出图;先生成 / 读取 template_profile.json,再按“Codex 原生路径”直接调当前会话的图片生成 tool 输出第 1 页 PNG,用户确认后再生成全量页面并打包。 5. 告知用户产物路径
模板页面挑选 / 复用原则
核心原则:尽量做到 1 page : 1 layout----同一份 deck 里每个 slide 用不同的模板页作 reference,观众会觉得每页都是新内容;如果同一个独特 layout 出现 2-3 次,观众下意识会想"为什么又是这页"。
vision 分析时会给每个 layout 标 reuse_friendly:
| reuse_friendly | 典型 layout | 多次使用的代价 |
|---|---|---|
false(不可复用,(!) 强警告) | 封面、3 个具名角色插画页、独特场景图(雪山/广播塔/复古收音机)、5 步骤 zigzag 各步独有图标、novelty 数据中央装置 | 视觉重复非常明显,观众会困惑 |
true(可复用,但仍建议错开,(i) 弱提示) | 纯文字、卡片网格、通用列表、章节小节标题 | 不致命,但平白浪费模板里的其它好版式 |
Agent 在搭 plan 时的执行策略: 1. 优先把模板里 N 个不同 layout 分配给 N 页 slide(N 不够就在 SKILL 里看 reuse_friendly=true 的部分挑能复用的) 2. 如果 plan 里某页内容结构非常相似(比如多个"5 步骤流程"),先尝试改写内容用不同 layout 表达(4 步骤 + 5 步骤分别用不同流程页),而不是同一个 zigzag 用两次 3. 冒烟跑完后,看 `Layout 复用检测` 那段输出:(!) 必须改,(i) 看情况改;改 plan 里相应 slide 的 layout_id 即可 4. 看完 profile JSON 选 layout:cat <cwd>/template_cache/<sha256>.json | jq '.layouts[] | {id, page_type, reuse_friendly, summary}';如果是多模态 agent 自己生成的 template_profile.json,就读取那个文件。
generate_ppt.py 在派发任务前会自动跑一次复用检测,把警告打到终端,不阻塞执行。
Skill 调用规范
当用户说"做一份 PPT" / "生成幻灯片"时:
1. 先问三件事(不要直接动手):
- 内容 / 页数 / 观众是谁?
- 风格偏好?按
styles/和docs/distilled-styles.md的场景类目映射推荐 1-2 个;或者用户上传自己的 .pptx 模板(走--template-pptx,自动渲染) - 是否需要单页测试一张图先看效果(API 直连用
--slides 1;Codex 原生路径直接生成第 1 页 PNG)
2. 先写 slides_plan.md 给用户确认文案(md 是 source of truth,人审阅友好) 3. 转 slides_plan.json:python3 scripts/md_to_plan.py slides_plan.md -o slides_plan.json(json 标为 generated,不手改;要改文案回到 md 改再转) 4. 构造 slide_spec(Agent 步骤):读 styles/<id>.md 了解视觉规范,然后为 slides_plan.json 每页构造 slide_spec,写入每页的 slide_spec 字段(格式见"指哪改哪"章节)。这一步让后续修改能精确到每个元素 5. 出图冒烟:
- API 直连 / 非 Codex 原生路径:跑
generate_ppt.py --slides 1出封面冒烟,效果 OK 再跑全量 - 当前 agent 就是带原生出图能力的 Codex:按上方“Codex 原生路径”直接调当前会话的
image_generationtool 生成outputs/<timestamp>/images/slide-01.png;不要用generate_ppt.py --backend codex
6. 告知用户产物路径,产物在 outputs/<timestamp>/,<title>.pptx 可直接打开
面向用户的表达规范
对普通用户汇报时,默认不要解释 slide_spec、metadata.json、element-updates、JSON Schema、pytest 命令等内部实现,除非用户明确问技术细节。
用户最关心的是:
1. 做出来的 PPT 是否好看、能不能直接用。 2. 哪些场景稳定,哪些场景需要人工验收。 3. 是否只改了指定页 / 指定内容。 4. 最终输出目录和 PPTX 文件在哪里。 5. 当前不足是什么,例如背景/文字仍是整页图片、数字小字需复核;真实 logo/产品图需提供素材,才能作为独立图片对象放入。
如果需要介绍修改能力,优先引用 docs/edit_guide.md 的场景分级、before / after 案例和当前不足,不要把实现链路放在用户前面。
当用户说"改第 X 页的 XX"时:
1. 找到 session:先用 --list-sessions 列出现有 session,确定要编辑的是哪个 2. 读 metadata.json:查看对应 slide 的 slide_spec,找到目标元素 3. 确认修改内容:告知用户当前内容,询问新内容 4. 构造编辑指令:用 --element-updates 指定要改的元素和新内容 5. 执行 --edit:python3 scripts/generate_ppt.py --edit X --session <ts> --element-updates '{"elem_id": {"content": "new"}}' 6. 告知结果:新版本已生成,PPTX 已更新
仅生成部分页
python3 scripts/generate_ppt.py --plan my_plan.json --style styles/dark-aurora.md --slides 1,3,5跑过的页有同名 PNG 时会自动跳过,方便逐页迭代。
"指哪改哪" PPT 编辑工作流
本节只保留 agent 执行所需信息;完整结构、版本链和数据安全细节见 docs/workflow.md。向普通用户解释能力时看 docs/edit_guide.md。
生成时的结构化要求
生成新 PPT 时,agent 应为每页写入简洁的 slide_spec,方便后续按标题、副标题、卡片、指标等元素精确修改。最低要求:
layout: 一句话描述版式elements: 语义化元素字典- 常用元素 ID:
title、subtitle、card_1、card_2、metric_1、date_line、footer - 元素字段优先写
type、content或heading/body、position、style、color
不要为了追求完整而写很长的 spec;能让后续定位和编辑即可。
修改已有幻灯片
1. 用 --list-sessions 找到目标 session。 2. 读目标 session 的 metadata.json,定位页号和元素 ID。 3. 如果用户没给新内容,先问一句“当前是 X,要改成什么?” 4. 用 --element-updates 描述修改;需要更强约束时加 --edit-instruction,明确“其他内容、布局、配色不要动”。 5. 执行 --edit 后检查目标页图片,重建后的 PPTX 会同步更新。
常用命令:
python3 scripts/generate_ppt.py \
--edit 3 \
--session 20240523_143052 \
--element-updates '{"subtitle": {"content": "医疗数据碎片化"}}' \
--edit-instruction "将副标题从'健康管理的两类割裂'改为'医疗数据碎片化'"
python3 scripts/generate_ppt.py \
--edit 3 \
--session 20240523_143052 \
--edit-prompt "在参考图基础上,只修改标题下方副标题的文字,从'健康管理的两类割裂'改为'医疗数据碎片化',保持位置、字体、颜色、大小完全不变。" \
--element-updates '{"subtitle": {"content": "医疗数据碎片化"}}'外部 PPTX 摄取后修改
外部 PPTX 可先摄取为图片 session,再由多模态 agent 看图补齐每页的 slide_spec。没有补 spec 前,不要承诺“精确改某个对象”。
python3 scripts/generate_ppt.py --ingest-pptx path/to/deck.pptx
python3 scripts/generate_ppt.py --ingest-pptx path/to/deck.pptx --session my_deck_2024回滚和 session 列表
python3 scripts/generate_ppt.py \
--rollback 3 \
--to-version 1 \
--session 20240523_143052
python3 scripts/generate_ppt.py --list-sessions文件结构
gpt-image2-ppt-skills/
|---- SKILL.md # 本文件(Claude Code skill 入口)
|---- AGENTS.md # codex / aider / cursor 等 agent 的薄索引,指向本文件
|---- README.md # 项目说明
|---- scripts/ # 所有 Python 脚本
| |---- generate_ppt.py # 主入口(CLI)
| |---- md_to_plan.py # slides_plan.md -> slides_plan.json 转换器(CLI)
| |---- render_template.py # PPTX -> 每页 PNG 的辅助脚本(CLI + library)
| |---- image_generator.py # gpt-image-2 wrapper(支持 reference image,openai backend)
| |---- codex_backend.py # 可选:走 codex CLI 出图(--backend codex)
| \---- template_analyzer.py # PPT 模板剖析器(vision + 缓存)
|---- styles/ # 所有可用风格,每个 .md 对应一个 style id
| |---- gradient-glass.md dark-aurora.md
| |---- clean-tech-blue.md risograph.md
| |---- vector-illustration.md japanese-wabi.md
| |---- editorial-mono.md swiss-grid.md
| |---- hand-sketch.md y2k-chrome.md
|---- docs/README.en.md # 英文 README
|---- install_as_skill.sh # 一键安装到 agent skills 目录
|---- requirements.txt # requests + python-dotenv + python-pptx + jsonschema + pymupdf
\---- .env.example调用时产生的运行时目录都在 <cwd> 下:
<your-project>/
|---- template_renders/<stem>/page-NN.png # PPTX 渲染(render_template.py)
|---- template_cache/<sha256>.json # 外挂 vision 风格分析缓存(纯文本 agent 路径)
|---- template_profile.json # 多模态 agent 可手写/生成的模板 profile(可选)
\---- outputs/<timestamp>/ # 每次生成产物License
Apache License 2.0.
# OpenAI official Images API (or any OpenAI-compatible relay you trust)
OPENAI_BASE_URL=https://api.openai.com
OPENAI_API_KEY=sk-your-key-here
# Model name (default: gpt-image-2)
GPT_IMAGE_MODEL_NAME=gpt-image-2
# Image quality: low / medium / high / auto
GPT_IMAGE_QUALITY=high
# -- Vision provider (OPTIONAL) ----------------------------------------------
# Only required when the caller agent is text-only and cannot inspect rendered
# .pptx template screenshots itself. Multimodal agents can write
# template_profile.json and pass it with --template-profile instead.
# Pick a provider you already trust -- there is no recommended endpoint default.
#
# Examples (replace with your own endpoint and key; never paste random URLs):
# VISION_BASE_URL=https://your-openai-compatible-relay.example.com/v1
# VISION_API_KEY=sk-your-vision-key-here
# VISION_MODEL_NAME=gemini-3.1-pro-preview # or any vision-capable model
#
# Leave these unset when a multimodal agent will analyze template screenshots
# itself, or when you do not use template-clone mode.
VISION_BASE_URL=
VISION_API_KEY=
VISION_MODEL_NAME=
* text=auto
*.sh text eol=lf
*.py text eol=lf
*.md text eol=lf
*.json text eol=lf
*.yaml text eol=lf
*.yml text eol=lf
*.html text eol=lf
*.css text eol=lf
*.js text eol=lf
# Byte-compiled / optimized / DLL files
__pycache__/
*.py[codz]
*$py.class
# C extensions
*.so
# Distribution / packaging
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
share/python-wheels/
*.egg-info/
.installed.cfg
*.egg
MANIFEST
# PyInstaller
# Usually these files are written by a python script from a template
# before PyInstaller builds the exe, so as to inject date/other infos into it.
*.manifest
*.spec
# Installer logs
pip-log.txt
pip-delete-this-directory.txt
# Unit test / coverage reports
htmlcov/
.tox/
.nox/
.coverage
.coverage.*
.cache
nosetests.xml
coverage.xml
*.cover
*.py.cover
.hypothesis/
.pytest_cache/
cover/
# Translations
*.mo
*.pot
# Django stuff:
*.log
local_settings.py
db.sqlite3
db.sqlite3-journal
# Flask stuff:
instance/
.webassets-cache
# Scrapy stuff:
.scrapy
# Sphinx documentation
docs/_build/
# PyBuilder
.pybuilder/
target/
# Jupyter Notebook
.ipynb_checkpoints
# IPython
profile_default/
ipython_config.py
# pyenv
# For a library or package, you might want to ignore these files since the code is
# intended to run in multiple environments; otherwise, check them in:
# .python-version
# pipenv
# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
# However, in case of collaboration, if having platform-specific dependencies or dependencies
# having no cross-platform support, pipenv may install dependencies that don't work, or not
# install all needed dependencies.
#Pipfile.lock
# UV
# Similar to Pipfile.lock, it is generally recommended to include uv.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
#uv.lock
# poetry
# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
#poetry.lock
#poetry.toml
# pdm
# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
# pdm recommends including project-wide configuration in pdm.toml, but excluding .pdm-python.
# https://pdm-project.org/en/latest/usage/project/#working-with-version-control
#pdm.lock
#pdm.toml
.pdm-python
.pdm-build/
# pixi
# Similar to Pipfile.lock, it is generally recommended to include pixi.lock in version control.
#pixi.lock
# Pixi creates a virtual environment in the .pixi directory, just like venv module creates one
# in the .venv directory. It is recommended not to include this directory in version control.
.pixi
# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
__pypackages__/
# Celery stuff
celerybeat-schedule
celerybeat.pid
# SageMath parsed files
*.sage.py
# Environments
.env
.envrc
.venv
env/
venv/
ENV/
env.bak/
venv.bak/
# Spyder project settings
.spyderproject
.spyproject
# Rope project settings
.ropeproject
# mkdocs documentation
/site
# mypy
.mypy_cache/
.dmypy.json
dmypy.json
# Pyre type checker
.pyre/
# pytype static type analyzer
.pytype/
# Cython debug symbols
cython_debug/
# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
# and can be added to the global gitignore or merged into this file. For a more nuclear
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/
# Abstra
# Abstra is an AI-powered process automation framework.
# Ignore directories containing user credentials, local state, and settings.
# Learn more at https://abstra.io/docs
.abstra/
# Visual Studio Code
# Visual Studio Code specific template is maintained in a separate VisualStudioCode.gitignore
# that can be found at https://github.com/github/gitignore/blob/main/Global/VisualStudioCode.gitignore
# and can be added to the global gitignore or merged into this file. However, if you prefer,
# you could uncomment the following to ignore the entire vscode folder
# .vscode/
# Ruff stuff:
.ruff_cache/
# PyPI configuration file
.pypirc
# Cursor
# Cursor is an AI-powered code editor. `.cursorignore` specifies files/directories to
# exclude from AI features like autocomplete and code analysis. Recommended for sensitive data
# refer to https://docs.cursor.com/context/ignore-files
.cursorignore
.cursorindexingignore
# Marimo
marimo/_static/
marimo/_lsp/
__marimo__/
# gpt-image2-ppt-skills
outputs/
template_cache/
template_renders/
tmp_*
docs/assets/gallery/
scripts/generate_style_gallery.py
.claude/
results/
logs/
outputs_cache/
output/
test_codex_output.png
test_runs/
# Local evaluation / test harnesses are useful while developing, but should not
# be published as part of the skill package.
tests/
evals/
benchmarks/
scripts/gen_demo_images.py
scripts/gen_edit_case_images.py
AGENTS.md -- 给 codex / aider / cursor 等 agent 的入口说明
本仓库是一个用 OpenAI gpt-image-2 生成 PPT 幻灯片的工具包,权威文档在 [`SKILL.md`](./SKILL.md)。任何涉及"做一份 PPT / 生成幻灯片 / 出图 / 改风格 / 按模板仿作"的请求,都先完整读 SKILL.md 再动手,不要凭本文件的摘要就开跑 -- 下面只是索引。
一分钟索引
- 主入口 CLI(API 直连 / 非 Codex 原生出图路径):
python3 scripts/generate_ppt.py --plan slides_plan.json --style styles/<id>.md - 内容源稿:先写
slides_plan.md(人审阅),再python3 scripts/md_to_plan.py slides_plan.md -o slides_plan.json;json 标为 derived,不手改 - 十种内置风格:见
styles/目录 +SKILL.md顶部表格 - 模板克隆:
--template-pptx path/to/xxx.pptx --template-strict,vision 分析 + 缓存细节在SKILL.md的"模板克隆模式"一节;如果你自己就是多模态 agent(多模态 Claude / GPT / 原生 Codex 等),可以直接Readtemplate_renders/<stem>/page-*.png自己抽风格写template_profile.json,用--template-profile传入,不用外挂VISION_* - 冒烟策略:API 直连 / 非 Codex 原生路径先
--slides 1出封面;如果你自己就是带原生出图能力的 Codex,则直接用当前会话的 image_generation tool 生成第 1 页 PNG 做冒烟,不要经--backend codex - 产物:
<cwd>/outputs/<timestamp>/{images/, prompts.json, metadata.json, <title>.pptx}
调用规范(对 agent 的硬约束)
1. 先问三件事:内容 / 观众、风格偏好(或是否带 .pptx 模板)、是否先单页冒烟 2. 永远先 md 后 json:用户改文案改 md,不手改 json 3. 先冒烟再全量:API 直连用 --slides 1;Codex 原生出图用当前会话 image_generation tool 先生成封面 PNG 4. 告知产物路径:跑完把 outputs/<timestamp>/ 和 .pptx 路径明确告诉用户
凭据 / backend
- 默认后端
openai,需通过 agent 配置 / 系统环境变量注入OPENAI_API_KEY;standalone CLI 可用$GPT_IMAGE2_PPT_ENV或 skill 安装目录.env作为 fallback,不要写进调用者业务项目根目录.env - 如果你就是 codex:走「原生 image_generation 出图」路径——直接用你自带的
image_generationtool 出图,不要跑generate_ppt.py --backend codex(那是给 Claude Code 等非 codex caller 用的,会 spawn 一个多余的 codex 子进程)。完整流程见SKILL.md的「如果你就是 Codex agent」一节 - 脚本不会向上递归读取调用者项目目录的
.env,避免误吃业务项目密钥
不要做的事
- 不要手改
slides_plan.json(改 md 再转) - 不要跳过
SKILL.md"模板页面挑选 / 复用原则"直接用同一个 layout 给多页 slide 当 reference - 不要在没有
OPENAI_API_KEY、不能走 Codex 原生出图、且没有本地 codex CLI 的情况下强跑 -- 先检查环境
同步提醒
本文件是薄索引。如果 SKILL.md 有新增章节或流程变更,只在 SKILL.md 里维护,本文件的锚点列表按需补;正文描述不要复制到这里,避免两边漂移。
interface:
display_name: "GPT-Image-2 PPT Generator"
short_description: "Create polished PPT decks from a topic or an existing .pptx template. Preview one slide first, then output high-res slide images and a 16:9 .pptx."
default_prompt: "Make me a 5-slide PPT about [topic]. Pick a fitting style and show me the cover first."
# Trust surface (what the user agrees to when installing this skill):
# - Reads process env, platform-injected env, explicit GPT_IMAGE2_PPT_ENV,
# and skill-owned .env fallback only; never walks parent project directories.
# - Hits exactly the OPENAI_BASE_URL endpoint you set, plus optional
# VISION_BASE_URL endpoint you set (only when the caller agent is text-only).
# - Downloads images returned by the model API; URLs are printed before
# download, capped at 50MB, http/https only.
# - Uses PowerPoint / Keynote / LibreOffice when available to render template .pptx -> PNG.
# - Writes outputs to <cwd>/{template_renders,template_cache,outputs}/.
扩展风格库
2026-05-26 从公开渠道 500+ 个 PPT 模板 中筛选补充了 22 个优质风格。这些风格已经写入 `styles/`。后续还会持续补充,也欢迎大家提供好看的 PPT 模板或风格参考。
风格列表
| 展示 | 视觉特色 | 适合场景 |
|---|---|---|
| <img src="assets/distilled-styles/abstract-art-showcase.jpg" width="640"><br><sub>abstract-art-showcase</sub> | 黑白极简、艺术展览感、超大字体和抽象画面并置 | 艺术策展、作品集、品牌调性展示 |
| <img src="assets/distilled-styles/coal-industry-business-company-profile.jpg" width="640"><br><sub>coal-industry-business-company-profile</sub> | 工业棕黑、粗重标题、结构线和硬朗图标 | 能源、制造业、重资产公司介绍 |
| <img src="assets/distilled-styles/college-candy-aesthetics-infographics.jpg" width="640"><br><sub>college-candy-aesthetics-infographics</sub> | 糖果色、校园感、圆润信息图和轻快装饰 | 教育、校园活动、轻量数据科普 |
| <img src="assets/distilled-styles/creative-agency.jpg" width="640"><br><sub>creative-agency</sub> | 创意机构气质、强视觉拼贴、鲜明版式节奏 | Agency 提案、品牌方案、创意汇报 |
| <img src="assets/distilled-styles/culinary-innovation.jpg" width="640"><br><sub>culinary-innovation</sub> | 餐饮创新感、食材摄影、暖色块和杂志式排版 | 餐饮品牌、食品创新、菜单/新品发布 |
| <img src="assets/distilled-styles/data-science-consulting.jpg" width="640"><br><sub>data-science-consulting</sub> | 数据咨询蓝灰、模块化布局、图表和技术感信息层级 | 数据分析、AI 咨询、企业数字化 |
| <img src="assets/distilled-styles/mindfulness-in-the-classroom-breathing-techniques.jpg" width="640"><br><sub>mindfulness-in-the-classroom-breathing-techniques</sub> | 柔和心理健康配色、留白、圆角块和安静插画感 | 心理健康、课堂活动、呼吸训练课程 |
| <img src="assets/distilled-styles/mind-maps-workshop-professional.jpg" width="640"><br><sub>mind-maps-workshop-professional</sub> | 专业工作坊风、思维导图节点、清晰流程结构 | 培训工作坊、方法论、团队共创 |
| <img src="assets/distilled-styles/meeting-agenda.jpg" width="640"><br><sub>meeting-agenda</sub> | 会议议程感、干净网格、强信息分组和商务标题 | 例会、项目同步、管理层汇报 |
| <img src="assets/distilled-styles/investment-company-business-plan.jpg" width="640"><br><sub>investment-company-business-plan</sub> | 投资机构质感、深浅对比、稳重商务版式 | 投资计划、基金介绍、商业计划书 |
| <img src="assets/distilled-styles/indigenous-cultures.jpg" width="640"><br><sub>indigenous-cultures</sub> | 文化纹样、自然色、手工质感和叙事型构图 | 文化课程、历史主题、公益教育 |
| <img src="assets/distilled-styles/health-disparities-and-social-determinants-of-health-doctor-of-philosophy-phd-in-health-behavior-and-health-education.jpg" width="640"><br><sub>health-disparities-and-social-determinants-of-health-doctor-of-philosophy-phd-in-health-behavior-and-health-education</sub> | 公共健康学术风、理性网格、柔和医疗色和论文感层级 | 医学论文答辩、公共健康报告、教育研究 |
| <img src="assets/distilled-styles/geometric-duotone-thesis.jpg" width="640"><br><sub>geometric-duotone-thesis</sub> | 双色几何、论文答辩感、斜切图形和强标题 | 学术答辩、研究报告、章节型内容 |
| <img src="assets/distilled-styles/geometric-clinical-case.jpg" width="640"><br><sub>geometric-clinical-case</sub> | 几何医疗风、冷静配色、病例卡片和清晰分栏 | 临床病例、医疗培训、诊疗汇报 |
| <img src="assets/distilled-styles/geometric-business.jpg" width="640"><br><sub>geometric-business</sub> | 商务几何块、稳健蓝绿调、简洁图表语言 | 商业计划、团队汇报、产品策略 |
| <img src="assets/distilled-styles/formal-lavender-portfolio.jpg" width="640"><br><sub>formal-lavender-portfolio</sub> | 淡紫正式感、作品集留白、优雅细线和柔和版式 | 个人作品集、设计简历、专业展示 |
| <img src="assets/distilled-styles/flowery.jpg" width="640"><br><sub>flowery</sub> | 花卉装饰、柔和色块、浪漫但有秩序的排版 | 生活方式、女性品牌、活动介绍 |
| <img src="assets/distilled-styles/first-impressions.jpg" width="640"><br><sub>first-impressions</sub> | 第一印象主题、强封面视觉、人物/标题的戏剧化关系 | 面试培训、个人品牌、沟通课程 |
| <img src="assets/distilled-styles/final-year-project-thesis-defense.jpg" width="640"><br><sub>final-year-project-thesis-defense</sub> | 毕业设计答辩、学院派网格、清晰章节与数据页 | 毕业答辩、项目结题、研究展示 |
| <img src="assets/distilled-styles/fashion-business-consulting-toolkit-aesthetic.jpg" width="640"><br><sub>fashion-business-consulting-toolkit-aesthetic</sub> | 时尚咨询感、高级拼贴、杂志排版和中性色 | 时尚商业、品牌咨询、趋势报告 |
| <img src="assets/distilled-styles/economic-impact-of-coronavirus.jpg" width="640"><br><sub>economic-impact-of-coronavirus</sub> | 经济影响报告风、严肃信息图、冷静色彩和数据叙事 | 宏观经济、政策分析、风险报告 |
| <img src="assets/distilled-styles/eco-green-business-plan.jpg" width="640"><br><sub>eco-green-business-plan</sub> | 鼠尾草绿、自然材质摄影、环保商务与极简分屏 | 可持续商业、环保品牌、健康生活方式 |
gpt-image2-ppt:PPT 修改能力测评与使用建议
一句话结论
gpt-image2-ppt 适合用自然语言修改 PPT 的视觉结果,例如改标题、换副标题、更新日期、删页脚、改数据卡片、给某页加小标识。它会把目标页重新生成成一张高质量 16:9 图片,再打包进 PPTX。
需要说清楚的是:它不是 PowerPoint 原生对象编辑工具。用户不能指望像手动点选文本框那样做到 100% 对象级、像素级不变;重要交付前仍然要人工看一遍。
---
用户最关心什么
| 用户关心的问题 | 当前回答 |
|---|---|
| 能不能直接说人话改 PPT? | 可以,例如“把第 3 页标题改成年度战略复盘,其他不要动”。 |
| 常见文字修改稳不稳? | 标题、副标题、日期、页脚这类短文本最稳定。 |
| 能不能只改某一页? | 可以,复杂多页 PPT 里可以只更新指定页,其他页不重新生成。 |
| 能不能改数据页? | 可以改多个指标,但数字类页面必须逐项复核。 |
| 能不能改别人的 PPT 模板? | 可以先导入或仿模板,但越要求像素级一致,越需要人工验收。 |
| 交付物是什么? | 每页高清 PNG、16:9 PPTX,以及生成过程记录。 |
---
场景能力总览
| 场景 | 稳定性 | 建议 |
|---|---|---|
| 改标题 / 副标题 / 日期 / 地点 | 高 | 适合直接使用,也是最推荐的修改场景。 |
| 同时改多处短文本 | 高 | 一次说清楚所有改动,并补一句“其他不要动”。 |
| 删除页脚、小标签、小文本 | 中高 | 通常可用,但要看删除区域是否有轻微重绘痕迹。 |
| 更新数据卡片和关键数字 | 中高 | 可用,但必须逐项核对数字、单位和位置。 |
| 新增 logo / 图标 / 小标识 | 中 | 适合生成风格化小标识;真实品牌 logo 需要提供明确素材。 |
| 模板克隆后局部修改 | 中 | 可用,但要检查是否偏离原模板风格。 |
| 密集表格、财务报表、合同长文 | 低 | 不建议直接承诺,容易出现小字和数字误差。 |
| 原生 PPT 对象级编辑 | 部分支持 | 背景与文字仍是整页图片;通过 external_image 放入的真实图片会作为独立 PPT 图片对象叠加,可单独选中拖动。 |
---
场景 1:只修改封面标题
用户需求:
把标题改成「年度战略复盘」,其他所有内容、布局、配色、装饰都不要动。| 修改前 | 修改后 |
|---|---|
| <img src="assets/demo1_before.jpg" width="100%" alt="修改前:产品发布会"> | <img src="assets/demo1_after.jpg" width="100%" alt="修改后:年度战略复盘"> |
测评结果:
| 检查项 | 结果 |
|---|---|
| 标题更新 | 通过,主标题已替换。 |
| 副标题保持 | 通过,副标题未被误改。 |
| 布局保持 | 通过,居中排版和视觉重心保持。 |
| 背景装饰保持 | 通过,整体风格一致。 |
结论:单个短文本替换是最稳定的场景,适合对外演示和日常交付。
---
场景 2:同时修改标题和底部日期
用户需求:
标题改成「2025 年度战略发布会」,
底部日期改成「2025年3月 · 深圳」,
其他都不要动。| 修改前 | 修改后 |
|---|---|
| <img src="assets/demo2_before.jpg" width="100%" alt="修改前:2024 年度回顾"> | <img src="assets/demo2_after.jpg" width="100%" alt="修改后:2025 年度战略发布会"> |
测评结果:
| 检查项 | 结果 |
|---|---|
| 标题更新 | 通过。 |
| 日期地点更新 | 通过。 |
| 副标题保持 | 通过。 |
| 背景和构图保持 | 通过。 |
结论:一次修改多个清晰文本元素可行。用户最好把要改的内容列完整,并说明哪些内容不能动。
---
场景 3:删除底部版权文字
用户需求:
底部那行 Copyright 字样去掉,其他都不要动。| 修改前 | 修改后 |
|---|---|
| <img src="assets/demo3_before.jpg" width="100%" alt="修改前:带版权文字"> | <img src="assets/demo3_after.jpg" width="100%" alt="修改后:删除版权文字"> |
测评结果:
| 检查项 | 结果 |
|---|---|
| 版权文字删除 | 通过。 |
| 标题和副标题保持 | 通过。 |
| 背景保持 | 通过。 |
| 主体构图保持 | 通过。 |
结论:删除小文本可用,但比单纯替换文字略有风险,因为删除区域需要重新补背景。
---
场景 4:批量更新数据指标
用户需求:
只更新三张数据卡片里的文字:
左侧改成「用户增长 520%」,
中间改成「营收突破 3.6亿」,
右侧改成「满意度 99.5%」。
标题、卡片位置、颜色、背景、装饰都不要动。| 修改前 | 修改后 |
|---|---|
| <img src="assets/demo4_data_before.jpg" width="100%" alt="修改前:三列数据指标"> | <img src="assets/demo4_data_after.jpg" width="100%" alt="修改后:三列数据指标更新"> |
测评结果:
| 检查项 | 结果 |
|---|---|
| 左侧指标更新 | 通过。 |
| 中间指标更新 | 通过。 |
| 右侧指标更新 | 通过。 |
| 标题保持 | 通过。 |
| 三列卡片结构保持 | 通过。 |
结论:数据页可以批量改,但数据是用户最敏感的部分。正式交付前必须逐项核对数字、单位、小数点和百分号。
---
场景 5:右上角新增 logo
用户需求:
在右上角增加一个简洁的科技公司 logo 图标,
大小约占页面宽度的 7%,不要遮挡标题和副标题。
其他文字、背景、布局、配色、装饰都不要动。| 修改前 | 修改后 |
|---|---|
| <img src="assets/demo5_logo_before.jpg" width="100%" alt="修改前:无 logo"> | <img src="assets/demo5_logo_after.jpg" width="100%" alt="修改后:右上角新增 logo"> |
测评结果:
| 检查项 | 结果 |
|---|---|
| logo 新增 | 通过,右上角出现小标识。 |
| 不遮挡主文案 | 通过。 |
| 标题和副标题保持 | 通过。 |
| 背景构图保持 | 通过。 |
结论:新增小型视觉元素可用。注意:如果用户需要真实公司 logo,不应该只靠文字描述生成,而应提供 logo 素材。
---
复杂多页:只修改某一页
用户经常会问:“我有一整套 PPT,只想改第 7 页,会不会影响其他页?”
当前结论:可以只更新指定页。下面的多页证据图展示了修改前后只有第 2 页发生变化,第 1 页和第 3 页保持原样。
<p align="center"> <img src="assets/demo6_multislide_evidence.jpg" width="100%" alt="复杂多页只改目标页证据图"> </p>
| 检查项 | 结果 |
|---|---|
| 目标页更新 | 通过。 |
| 其他页保持 | 通过。 |
| 页序保持 | 通过。 |
| 适合复杂多页 PPT | 适合,但目标页本身仍需人工验收。 |
---
用户可以怎么说
用户不需要懂命令、JSON 或内部结构。直接这样说就够了:
| 用户说法 | 适合程度 |
|---|---|
| “把封面标题改成年度战略复盘,其他不要动。” | 很适合 |
| “第 2 页底部日期改成 2025 年 3 月,地点改成深圳。” | 很适合 |
| “把第 5 页三个数据改成这三个新数字。” | 适合,但要核对数字 |
| “删掉每页底部版权文字。” | 适合,建议逐页检查 |
| “右上角加一个我们公司 logo。” | 需要提供 logo 图片素材 |
| “整套完全按原品牌手册像素级复刻。” | 不建议直接承诺 |
---
当前不足
| 不足 | 对用户的影响 | 建议说法 |
|---|---|---|
| 背景与文字是整页图片 | 在 PowerPoint 里不能直接点选文字框继续编辑;通过 external_image 放入的真实图片可单独拖动。 | “适合直接展示和分享;如果要原生可编辑文字框,需要后续增强;真实素材图可以作为独立对象放入。” |
| 生成式编辑可能有轻微漂移 | 背景纹理、字体细节、图标形态可能和原图不完全一致。 | “我会尽量保持不变,但交付前需要看一遍效果。” |
| 数字和小字需要复核 | 数据页、表格页、密集文字页对准确性要求高。 | “我可以改,但你需要重点检查数字和小字。” |
| 外部 PPTX 初次导入不等于完全可精确编辑 | 需要先让 AI 看图理解每页有哪些元素。 | “可以改,但复杂模板建议先做一页测试。” |
| 真实 logo / 品牌资产不能只靠描述 | 模型会生成相似风格的小图标,不一定是真实 logo。 | “请提供 logo 文件,我再放入或参考它生成。” |
| 像素级品牌一致性不稳定 | 对严格企业模板、法务文件、财报不适合作无条件承诺。 | “这类场景需要模板测试和人工验收。” |
---
交付前检查清单
| 检查项 | 为什么重要 |
|---|---|
| 标题、副标题是否正确 | 这是最显眼的错误来源。 |
| 数字、单位、百分号、小数点是否正确 | 数据页最容易被用户追责。 |
| 有没有误改其他文字 | 生成式编辑可能影响相邻区域。 |
| logo、图标是否符合品牌要求 | 真实品牌资产需要格外检查。 |
| PPTX 页面比例是否为 16:9 | 避免投影或会议屏幕变形。 |
| 是否只修改了用户指定页 | 多页交付前快速翻一遍。 |
---
最终总结
gpt-image2-ppt 已经适合覆盖大多数“让 AI 帮我改 PPT 视觉结果”的需求,尤其是标题、日期、短文案、页脚、数据卡片、局部小元素这类明确修改。
当前不适合承诺的是:PowerPoint 原生对象级编辑、复杂表格逐字零误差、严格品牌模板像素级一致、未提供真实素材却要求真实 logo。
最稳的使用方式是:先让用户用自然语言说清楚要改哪一页、改什么、哪些不动;先测一页确认风格,再批量处理;交付前按检查清单快速验收。
gpt-image2-ppt external real-image overlay logic
================================================
Goal:
Generate the PPT background with gpt-image-2, then insert real images as
independent PPT picture objects without covering generated text/graphics.
Core rule:
The skeleton reference image is not the planner.
Slot planning must happen before the skeleton is drawn.
+---------------------------+ +---------------------------+
| slide content profile | | real asset profile |
|---------------------------| |---------------------------|
| title length | | image count |
| body text load | | width / height ratio |
| page type | | photo / chart / document |
| template media slots | | high-detail / critical |
| layout/style tendency | | |
+-------------+-------------+ +-------------+-------------+
| |
+----------------+-------------------+
|
v
+-------------+-------------+
| slot planning |
|---------------------------|
| read template candidates |
| score candidate regions |
| preserve template rhythm |
| choose text-safe region |
| enlarge high-detail image |
| compact text if needed |
| write planning metadata |
| |
| output: |
| - position |
| - computed_bbox |
| - auto_layout_reason |
| - layout_planning_profile |
+-------------+-------------+
|
v
+-------------+-------------+
| transparent skeleton |
|---------------------------|
| draw only corner markers |
| no white canvas |
| no filled rectangle |
| no mask_placeholder shape |
| no visible placeholder |
+-------------+-------------+
|
v
+--------------------+--------------------+
| prompt construction |
|-----------------------------------------|
| remove conflicting fake-photo/template |
| directives |
| tell model to keep key content outside |
| corner-marker area |
| high-detail page: left text rail + |
| template-aware large asset panel |
+--------------------+--------------------+
|
v
+-------------+-------------+
| gpt-image-2 generation |
|---------------------------|
| generates background only |
| no real image |
| no fake image |
| no placeholder card |
+-------------+-------------+
|
v
+-------------+-------------+
| deterministic overlay |
|---------------------------|
| insert real image by code |
| same computed_bbox |
| independent PPT picture |
| user can drag/edit image |
+-------------+-------------+
|
v
+-------------+-------------+
| multimodal overlay review |
|---------------------------|
| compare background vs |
| final overlay preview |
| |
| check: |
| - text/graphic coverage |
| - white placeholder issue |
| - high-detail readability |
+-------------+-------------+
|
+----------+----------+
| |
v v
+--------+--------+ +-------+---------+
| pass | | fail |
|-----------------| |-----------------|
| write PPTX | | choose suggested|
| report success | | or fallback slot|
+-----------------+ | regenerate page |
| max 3 rounds |
+-------+---------+
|
v
+-------+---------+
| still fail |
|-----------------|
| exit nonzero |
| do not claim OK |
+-----------------+
Planning policy summary
-----------------------
1. Photos can be moderate-size visual anchors.
2. Charts/documents need larger readable panels.
3. Template/profile image slots are preferred over generic right/bottom slots.
If a template says this page is left-image/right-text, centered visual, or
bottom chart, the planner uses that composition first.
4. Dense diagrams / architecture images are critical assets. If the template
slot is too small for readability, the planner can reject it and fall back to
a larger readable region.
5. Multi-image pages use grids or paired panels after preserving the text area.
6. Long titles on high-detail pages are wrapped into short lines, and body text
is compacted to short bullets.
7. The transparent skeleton is created from the final computed_bbox so the
reference, generated background, overlay preview, and PPTX object all agree.
gpt-image2-ppt-skills 安装指南(给 AI agent 读)
本文档是让 AI 助手自动完成安装的可执行指引。Claude Code / OpenClaw / Codex / Cursor / Trae / Hermes Agent 等任一 agent 按以下步骤执行即可。
>
人类用户请不要手动照抄下面的步骤,直接把这份 URL 扔给你的 AI 助手,它自己会跑通。
项目一句话介绍
gpt-image2-ppt-skills 是一个 Claude Code / Codex / OpenClaw Skill,用 OpenAI gpt-image-2 生成视觉强烈的 PPT(10 套内置风格 + 支持仿任意 .pptx 模板),产出高清逐页 PNG + 16:9 .pptx。
仓库:https://github.com/JuneYaooo/gpt-image2-ppt-skills
前置依赖(agent 请先检查)
- 必需:
git、python3(3.8+)、pip - 可选:本机可执行的 PPTX 渲染后端(Windows PowerPoint / macOS Keynote / LibreOffice,仅模板克隆模式需要,用来把
.pptx渲染成 PNG)
如果缺 git / python3,先用系统包管理器装好;PPTX 渲染后端可以留到用户要用模板克隆模式时再装。鸿蒙 / Termux / 容器 / 特殊架构不要只看 which libreoffice,必须用 python3 scripts/render_template.py --check 确认二进制真的可运行。
安装步骤
1. 克隆仓库到临时目录
git clone https://github.com/JuneYaooo/gpt-image2-ppt-skills.git /tmp/gpt-image2-ppt-skills
cd /tmp/gpt-image2-ppt-skills2. 跑安装脚本
# Claude Code
bash install_as_skill.sh --target claude
# Codex
bash install_as_skill.sh --target codex这一步会:
- 把项目文件拷贝到对应 agent 的 skill 目录
pip install -r requirements.txt装齐 Python 依赖- 保留
.env.example;安装目录.env只作为 standalone CLI fallback
安装脚本是交互式的,如果目标目录已存在会问是否覆盖;agent 可以用 yes | bash install_as_skill.sh 自动选"是",或先检查目录是否存在再决定。
3. 只有在走 API 直连时才注入环境变量
如果当前 agent 就是带原生图片生成能力的 Codex,可以跳过这一步,直接重启后走 SKILL.md 里的原生出图路径。
否则,agent 需要主动问用户要 API key:
请提供你的 OpenAI API key(或任意 OpenAI 兼容中转的 base_url + key)。
优先通过当前 agent 框架或运行环境注入,不要写进调用者业务项目根目录 .env:
- Claude Code:用户级
~/.claude/settings.json,或项目级.claude/settings.local.json - OpenClaw / 自定义 Agent:用
apiKey/ env reference 引用系统环境变量,避免明文写进项目配置 - CI / 服务器:用系统环境变量、Docker Compose、Kubernetes Secret 或 CI Secret
- standalone CLI fallback:设置
GPT_IMAGE2_PPT_ENV=/path/to/private.env,或使用 skill 安装目录下的.env
需要注入的变量如下:
OPENAI_BASE_URL=https://api.openai.com # 或用户提供的兼容中转 URL
OPENAI_API_KEY=sk-... # 用户提供的 key(必需)
GPT_IMAGE_MODEL_NAME=gpt-image-2
GPT_IMAGE_QUALITY=high # low / medium / high / auto如果用户没有 OpenAI 官方 key:当前 agent 若是带原生出图能力的 Codex,优先走SKILL.md的 Codex 原生路径;否则可以告诉用户本 skill 支持--backend codex启动本地codex exec子进程复用登录态,或使用任意 OpenAI 兼容中转站(把OPENAI_BASE_URL改成中转站地址即可)。
3.1 模板克隆的视觉分析
模板克隆需要先看懂 .pptx 渲染出的页面截图:
- 当前 agent 是多模态模型(例如多模态 Claude / GPT / 原生 Codex)时,不需要配置
VISION_*;agent 直接读取template_renders/<stem>/page-*.png,生成template_profile.json,再用--template-profile传给 CLI。 - 当前 agent 是纯文本模型(例如 DeepSeek 文本模型)时,需要额外配置
VISION_BASE_URL/VISION_API_KEY/VISION_MODEL_NAME,由template_analyzer.py调独立多模态模型分析模板。
4. 提示用户重启 agent
装完之后,告诉用户:
已安装完成。请重启当前 agent(Claude Code / Codex / 其它)让 skill 生效。
5.(可选)清理临时目录
rm -rf /tmp/gpt-image2-ppt-skills冒烟测试(用户重启 agent 后)
告诉用户直接跟 agent 说:
帮我用 gpt-image2-ppt 生成一份关于「猫为什么是液体」的 3 页 PPT,风格用 gradient-glass。正常的话 agent 会自己写 slides_plan.md、转成 slides_plan.json,然后按当前环境分流:API 直连路径跑 scripts/generate_ppt.py --slides 1 先出封面;Codex 原生路径直接用当前会话的图片生成 tool 先出封面 PNG。确认后再跑全量,最后给出输出目录和 .pptx 的路径。
常见问题(给 agent 参考)
- `ModuleNotFoundError: pymupdf` → 在实际安装目录里重跑
pip install -r requirements.txt - `libreoffice: command not found` / `permission denied` / `Exec format error`(仅模板克隆模式)→ 先跑
python3 scripts/render_template.py --check;Linux 桌面可装apt install libreoffice,macOS 可装 Keynote 或brew install --cask libreoffice;鸿蒙 / Termux / 容器 / 特殊架构建议在桌面端把模板每页导出为page-01.png、page-02.png后用--template-images,不要依赖aspose-slides兜底 - `OPENAI_API_KEY 未设置` → 如果你不是走 Codex 原生路径,回到步骤 3,检查 agent / 系统环境变量是否已注入;standalone CLI 再检查
GPT_IMAGE2_PPT_ENV或 skill 安装目录.env - agent 识别不到 skill → 确认目录装到了对应 agent 的技能目录,并且完全重启过当前 agent
完成标志
以下三条都满足即视为安装成功:
1. 对应安装目录下的 SKILL.md 存在 2. 如果走 API 直连模式,agent / 系统环境变量中能提供可用 OPENAI_API_KEY 3. agent 重启后,用户用自然语言要求生成 PPT 时能触发本 skill
装完不用逐字读 SKILL.md,但需要告诉用户:"你可以直接用自然语言要 PPT,也可以把任意 .pptx 模板丢给我做克隆。"
PPT 实现流程图
┌────────────────────────────────────────────┐
│ 1. 用户输入 │
│ 主题 / 大纲 / 文案 / 图片 / Logo / PPT模板 │
└──────────────────────┬─────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ 2. 内容整理 │
│ - 明确观众和用途 │
│ - 确定页数 │
│ - 拆成封面、内容页、数据页、总结页等 │
│ - 标出每页标题、正文、数据、图片素材 │
└──────────────────────┬─────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ 3. 视觉方向选择 │
└──────────────┬─────────────────────┬───────┘
│ │
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ A. 使用内置风格 │ │ B. 使用用户 PPT 模板 │
│ 科技 / 商务 / 手绘等 │ │ 参考版式 / 配色 / 节奏│
└───────────┬──────────┘ └───────────┬──────────┘
│ │
└─────────────┬─────────────┘
▼
┌────────────────────────────────────────────┐
│ 4. 每页版式规划 │
│ - 标题放哪里 │
│ - 正文如何分组 │
│ - 数据如何突出 │
│ - 图片区域如何预留 │
│ - 模板页如何分配,尽量一页一个版式 │
└──────────────────────┬─────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ 5. 是否有用户提供的真实图片? │
└──────────────┬─────────────────────┬───────┘
│ │
│ 否 │ 是
▼ ▼
┌──────────────────────┐ ┌──────────────────────────────┐
│ 进入页面生成 │ │ 6. 判断真实图片的使用方式 │
└───────────┬──────────┘ └───────────┬────────────┬─────┘
│ │ │
│ │ 默认 │ 用户明确说:
│ │ 必须保真 │ 不必保真 / 可重绘
│ ▼ ▼
│ ┌──────────────────┐ ┌──────────────────┐
│ │ A. 原图保真后贴 │ │ B. 参考图融合重绘 │
│ │ │ │ │
│ │ 适合: │ │ 适合: │
│ │ Logo │ │ 氛围图 │
│ │ 产品截图 │ │ 场景图 │
│ │ 医疗影像 │ │ 非关键照片 │
│ │ 表格 / 图表 │ │ 风格参考图 │
│ │ 证据截图 │ │ │
│ │ 精确 UI │ │ 结果: │
│ │ │ │ 融合进整页画面 │
│ │ 结果: │ │ 不单独后贴 │
│ │ 原图独立贴入 PPT │ │ 不保证细节保真 │
│ └─────────┬────────┘ └────────┬─────────┘
│ │ │
└────────────────────────┴─────────┬─────────┘
▼
┌────────────────────────────────────────────┐
│ 7. 生成每一页画面 │
│ - 按 16:9 横版生成 │
│ - 保持整套视觉统一 │
│ - 中文字体和层级清晰 │
│ - 有原图后贴时,生成背景要避开图片区域 │
│ - 有参考图重绘时,把图片风格融合进页面 │
└──────────────────────┬─────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ 8. 生成中间产物 │
│ - 每页高清图片 │
│ - 每页生成记录 │
│ - 每页版本记录 │
│ - 真实图片位置记录 │
└──────────────────────┬─────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ 9. 打包成 PPT │
└──────────────┬─────────────────────┬───────┘
│ │
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ 普通页面 │ │ 有保真图片的页面 │
│ 整页高清图铺满幻灯片 │ │ 第1层:整页背景图 │
└───────────┬──────────┘ │ 第2层:真实图片对象 │
│ └───────────┬──────────┘
└─────────────┬─────────────┘
▼
┌────────────────────────────────────────────┐
│ 10. 质检 │
│ - 是否有生成失败页 │
│ - 是否保持 16:9 │
│ - 真实图是否遮挡文字 │
│ - 是否出现白色占位块 │
│ - 高精度图片是否需要人工核对 │
└──────────────────────┬─────────────────────┘
│
┌─────────┴─────────┐
│ │
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ 通过 │ │ 不通过 │
│ 交付 PPTX │ │ 调整版式 / 重生成该页 │
└───────────┬──────────┘ └───────────┬──────────┘
│ │
└─────────────┬─────────────┘
▼
┌────────────────────────────────────────────┐
│ 11. 后续修改 │
│ - 用户说改第几页 │
│ - 只重生成目标页 │
│ - 其他页保持不动 │
│ - 保留版本,可回滚 │
└────────────────────────────────────────────┘
关键规则:
1. 用户给真实图片时,默认走“原图保真后贴”。
2. 用户明确说不需要保真、追求融合效果时,才走“参考图融合重绘”。
3. 医疗影像、诊断图、表格、证据截图、论文图表、精确 UI 截图,不建议重绘。
4. 普通文字和背景主要是整页图片;保真图片可以作为独立对象放进 PPT。<div align="center">
gpt-image2-ppt-skills
Generate design-forward, highly polished PPT decks with OpenAI `gpt-image-2` in one shot.
Works natively in Claude Code, Codex, OpenClaw, Hermes, and any other Skill-compatible agent. Once installed in your agent, a single natural-language prompt yields 16:9 high-res images + a ready-to-send .pptx — or clones any reference .pptx template and reskins it with new content.
Possibly one of the best-looking AI PPT Skills available today. Instead of filling text into traditional templates, it uses the visual taste, composition, and layout strengths of gpt-image-2 to generate each slide as a complete visual composition, aiming for decks that look polished, consistent, and presentation-ready from cover to inner pages.
The project also includes dedicated optimization for editing image-based PPTs. You can describe the target slide and element in natural language, and the system regenerates that slide through image-to-image editing while trying to preserve the original style and layout. One important caveat: the background and text in these PPTs are full-slide images. If your workflow depends on manually editing native PowerPoint text boxes and individual objects, this may not be the right fit.
    
🌐 中文 → ../README.md
</div>
---
🎬 Demo: feed one template, get a fresh deck in that style
<table> <tr> <th width="50%">Input: any reference template (.pptx / image)</th> <th width="50%">Output: cloned layout + new content</th> </tr> <tr> <td><img src="assets/template-demo-input.jpg" width="100%" alt="input template"></td> <td><img src="assets/template-demo-output.jpg" width="100%" alt="generated output"></td> </tr> <tr> <td align="center"><sub>English infographic template (Mass Media Infographics)</sub></td> <td align="center"><sub>Same layout / palette / illustration vocabulary, content swapped to "how normal people make AI-powered social content"</sub></td> </tr> </table>
---
✨ What it does
- 🎨 10 curated styles + an expanded style library — built-ins include Spatial Glass / Tech Blue / Editorial Mono / Dark Aurora / Risograph / Wabi / Swiss Grid / Hand Sketch / Y2K Chrome / Vector Illustration; on 2026-05-26, 22 additional high-quality styles were selected from 500+ publicly available PPT templates
- 🪄 Template-clone mode — drop in any
.pptx; the agent follows its layout, palette, and illustration language, then swaps in your new content - 🎯 Precise natural-language edits — say "change slide 3's subtitle", "remove the footer", or "replace these three metrics"; the agent regenerates only the target slide through image-to-image editing while trying to preserve the original style and layout
- 🎮 Dual output — high-res PNG per slide + 16:9
.pptxready to use - ⚡ 10-way concurrency by default — a 10-page deck finishes in ~30s
- 🧪 Preview one slide first — approve the cover before generating the full deck
- 🧾 Trackable edits — changed slides and generated versions can be traced and rolled back
✅ Best-fit use cases
| Use case | Fit | Notes |
|---|---|---|
| Generate a new deck from a topic | Strong | Good for reports, pitches, training, courses, product intros. |
| Create a new deck from a company template | Strong | Provide a .pptx, approve one cover first, then run the full deck. |
| Edit titles, subtitles, dates, footers | Strong | The most stable editing scenario. |
| Update metric cards and key numbers | Good | Works, but every number must be checked before delivery. |
| Modify only one slide in a multi-slide deck | Good | The target slide is regenerated; other slides are left alone. |
| Dense tables, financial reports, legal long copy | Weak | Small text and numbers need strict human review. |
🎨 The 10 built-in styles
Below: the 10 styles each generating one cover under the same topic — "How to make a PPT with gpt-image-2". All covers are raw gpt-image-2 output, no PS.!10 styles · same topic, raw gpt-image-2 output
| Style ID | One-liner | Use cases |
|---|---|---|
gradient-glass | Apple Vision OS / Spatial Glass | AI product launches, technical talks, creative pitches |
clean-tech-blue | Stripe / Linear-grade blue & white | Investor decks, business plans, corporate strategy |
vector-illustration | Retro vector + black outlines | Education, brand storytelling, community sharing |
editorial-mono | Kinfolk / Monocle editorial | Brand reveals, cultural interviews, book talks |
dark-aurora | Linear / Vercel dark neon | AI products, dev tools, technical talks |
risograph | Riso 2-spot-color print + halftone | Creative studios, indie zines, design agencies |
japanese-wabi | Muji / Hara Kenya wabi-sabi | Tea ceremony, lifestyle, luxury, cultural lectures |
swiss-grid | Bauhaus / Vignelli international grid | Academic reports, museum exhibits, serious dashboards |
hand-sketch | Sketchnote / whiteboard | Workshops, product brainstorming, training |
y2k-chrome | Y2K liquid chrome + butterfly stickers | Streetwear, entertainment, brand collabs, Gen-Z marketing |
🧬 Expanded style library: 22 new styles added on 2026-05-26
On 2026-05-26, we added 22 high-quality styles selected from 500+ publicly available PPT templates. More styles will continue to be added, and good PPT template or style references are welcome.
See the full style table, thumbnails, style IDs, visual traits, and use cases in `distilled-styles.md`.
---
🧪 Editing Capability Report
If you care about "how reliable are edits in real scenes", see the user-facing case report:
- [`docs/edit_guide.md`](./edit_guide.md) — title replacement, date edits, footer removal, metric updates, logo insertion, single-slide edits in a multi-slide deck, current limitations, and a delivery checklist
Summary:
| Capability | Current behavior |
|---|---|
| Short text edits | Stable for everyday delivery. |
| Multiple explicit edits | Works best when the user clearly says what should stay unchanged. |
| Metric slides | Works, but numbers must be checked. |
| Small icon / logo insertion | Works for style-matched icons; real brand logos need source assets. |
| Native PowerPoint object editing | Not supported; output PPTX uses full-slide images. |
<details> <summary>Developer note: internal editing mechanism diagram</summary>
<img src="assets/architecture_cn.jpg" width="100%" alt="system architecture">
</details>
---
🚀 Install
Option 1: let your AI install it (recommended)
Paste this prompt into your AI assistant (Claude Code / OpenClaw / Codex / Cursor / Trae / Hermes Agent, or any other agent that supports Skills) and it will handle the install:
Please install gpt-image2-ppt-skills for me:
https://raw.githubusercontent.com/JuneYaooo/gpt-image2-ppt-skills/main/docs/install.mdThe agent will clone the repo, run the install script, ask for an API key only when direct API mode is needed, and tell you to restart.
Option 2: manual install
git clone git@github.com:JuneYaooo/gpt-image2-ppt-skills.git
cd gpt-image2-ppt-skills
bash install_as_skill.sh --target claude # Claude Code
# or
bash install_as_skill.sh --target codex # CodexThe script installs the skill into the selected agent directory:
- Claude Code:
~/.claude/skills/gpt-image2-ppt-skills/ - Codex:
~/.codex/skills/gpt-image2-ppt-skills/
If you use direct API mode, inject environment variables through your agent framework instead of writing secrets into the caller project's root .env:
- Claude Code: user-level
~/.claude/settings.json, or project-level.claude/settings.local.json - OpenClaw / custom agents: reference system env vars from
apiKey/ env config - CI / servers: system env vars, Docker Compose, Kubernetes Secrets, or CI Secrets
- Standalone CLI: set
GPT_IMAGE2_PPT_ENV=/path/to/private.env, or use the skill install directory.envas a fallback
# Variable names:
OPENAI_BASE_URL=https://api.openai.com # or any OpenAI-compatible relay
OPENAI_API_KEY=sk-... # required
GPT_IMAGE_MODEL_NAME=gpt-image-2
GPT_IMAGE_QUALITY=high # low / medium / high / autoIn Codex, if the current agent has native image generation, use the native path inSKILL.mdand skipOPENAI_API_KEY.
>
🔒 Won't accidentally eat your secrets: the script only reads the current process env, platform-injected variables, an explicitGPT_IMAGE2_PPT_ENV, and the skill install directory.envfallback. It does not walk up into caller project directories.
>
🪄 Template-clone mode additionally needs an executable PPTX renderer: Windows PowerPoint, macOS Keynote, or LibreOffice. Run python3 scripts/render_template.py --check first; HarmonyOS / Termux / containers / unusual architectures should not assume Linux aarch64 LibreOffice binaries are runnable.Vision analysis for template clone (optional)
In template-clone mode, the skill needs to "see" your .pptx template's visual style first. If your AI assistant is already multimodal (Claude Code with Claude Opus/Sonnet, Codex with GPT multimodal, etc.), the agent will analyze the visual style directly and generate a template_profile.json with reference_image to pass to the CLI with --template-profile. No extra configuration needed.
Only when your agent uses a text-only model (e.g., DeepSeek text model), you'll need the following env vars to use a separate multimodal model for template analysis:
# Optional: vision analysis for template clone (only needed by text-only agents; skip for multimodal agents)
VISION_BASE_URL=https://your-openai-compatible-relay.example.com/v1
VISION_API_KEY=sk-...
VISION_MODEL_NAME=gemini-3.1-pro-preview # or gpt-4o / claude-3.5-sonnet, any multimodal SKUSupports any multimodal model compatible with the OpenAI/v1/chat/completionsformat (Gemini / GPT-4o / Claude, etc.). Fully decoupled fromgpt-image-2— switching the vision provider won't affect image generation.
---
🛠 How to use inside Claude Code
Once installed, just say it in plain English:
Use gpt-image2-ppt to make a 5-slide deck about [your topic], style = dark-aurora.Template clone, same shape:
I have a company-template.pptx — make a 5-slide deck about [your topic] using that template.Claude will write the slides_plan, generate a cover first for you to approve, then run the full deck and hand back the .pptx path.
Prefer to call the CLI yourself instead of going through an agent? See `SKILL.md` — CLI flags and file layout live there.
---
🙏 Acknowledgements
- op7418/NanoBanana-PPT-Skills — reference for the original style prompts and early skill structure. This project swaps the image backend from Nano Banana Pro to OpenAI gpt-image-2, rewrites the 3 inherited styles and adds 7 new ones (10 total), and layers on template-clone mode (vision-based style extraction from any user
.pptx), an md-first authoring flow, automatic.pptxpackaging, and a codex CLI fallback backend. - lewislulu/html-ppt-skill — reference for the Claude Code skill
SKILL.mdfrontmatter.
💬 Community
**LINUX DO — Chinese Developer Community**
⭐ Star History

---
License
Apache License 2.0 — see LICENSE.
PPT 生成与编辑工作流
本文档说明 generate_ppt.py 的完整逻辑——从生成、编辑、回滚到外部 PPTX 摄取。
---
一、CLI 入口分发
generate_ppt.py
│
├── --list-sessions ──→ 扫描 outputs/ 下含 metadata.json 的目录,输出列表
│
├── --ingest-pptx ────→ 渲染外部 PPTX → 创建 session → 写入占位 metadata.json
│ (Agent 后续 Read PNG 填 slide_spec)
│
├── --edit ───────────→ cmd_edit_slide() ← 需 --session
│
├── --rollback ───────→ cmd_rollback_slide() ← 需 --session + --to-version
│
└── (默认) ───────────→ 生成模式 ← 需 --plan + (--style | --template-pptx)---
二、生成流程
2.1 内置风格 (--plan + --style)
slides_plan.md styles/<id>.md
│ │
│ md_to_plan.py │
▼ │
slides_plan.json │
│ │
│ Agent 综合两者,为每页构造 slide_spec │
│ (type / content / position / │
│ style / color) │
│ 写入每页 .slide_spec 字段 │
▼ ▼
┌────────────────────────────────────────────────────┐
│ generate_ppt.py │
│ │
│ slide_spec 有 elements? │
│ ├── YES → generate_prompt_from_spec() │
│ │ 逐元素描述位置、内容、样式 → 精确 prompt │
│ └── NO → generate_prompt() │
│ 自由格式 content → 基础 prompt │
│ │
│ 并发派发 → gpt-image-2 出图 │
│ │ │
│ ▼ │
│ outputs/<ts>/images/slide-NN.png │
│ │ │
│ ├── prompts.json (兼容旧格式) │
│ ├── metadata.json (slide_spec 版本历史) │
│ └── <title>.pptx (16:9 打包) │
└────────────────────────────────────────────────────┘2.2 模板克隆 (--template-pptx)
template.pptx
│
│ render_template.py (LibreOffice / Keynote / PowerPoint COM)
▼
template_renders/<stem>/page-NN.png
│
│ template_analyzer.py (vision 分析, 结果缓存)
▼
template_cache/<sha256>.json (TemplateProfile: layouts + global_style)
│
│ match_layout() → coerce_fields() → render_prompt_from_template()
▼
每页 prompt ──→ gpt-image-2
(可选 --template-strict → 模板页作 reference image)2.3 增量生成 (--slides + --output)
当 --output 指向已有 session 时,不会覆盖已有 metadata,而是合并新页到现有 session:
outputs/20240523_143052/ ← 已有 slide-01, slide-03
│
│ --output outputs/20240523_143052 --slides 2,4
▼
outputs/20240523_143052/ ← 新增 slide-02, slide-04
metadata.json 合并
prompts.json 重建
.pptx 更新---
三、编辑流程 (--edit)
用户: "改第 3 页的副标题"
│
▼
┌─ Agent ───────────────────────────────────────────────┐
│ 1. 读 metadata.json │
│ → slides["3"].versions[current_version] │
│ → spec.elements.subtitle │
│ content = "健康管理的两类割裂" │
│ position = "标题下方" │
│ │
│ 2. 问用户新内容 → "医疗数据碎片化" │
│ │
│ 3. 构造 CLI 参数: │
│ --element-updates │
│ '{"subtitle":{"content":"医疗数据碎片化"}}' │
│ --edit-instruction │
│ "将副标题从X改为Y" │
└───────────────────────────────────────────────────────┘
│
▼
┌─ cmd_edit_slide() ────────────────────────────────────┐
│ │
│ 1. 加载 metadata.json │
│ 2. 获取当前 slide_spec (latest version) │
│ 3. apply_spec_updates() → updated_spec │
│ 4. construct_edit_prompt(old_spec, updates): │
│ "在参考图基础上,只修改标题下方subheading的文字 │
│ 从「健康管理的两类割裂」改为「医疗数据碎片化」 │
│ 保持其他所有元素不变" │
│ │
│ 5. 备份原图: slide-03.png → slide-03_v0001.png │
│ 6. gpt-image-2 (edit_prompt + 原图 reference) → 新图 │
│ 7. 新图 → slide-03.png (覆盖) │
│ 8. 新图备份: slide-03.png → slide-03_v0002.png │
│ 9. metadata.json 追加新 version: │
│ { action:"edit", spec:<updated>, │
│ reference_version:2, edit_instruction:"..." } │
│ 10. 重建 .pptx │
└───────────────────────────────────────────────────────┘编辑 prompt 的两种方式
| 方式 | 参数 | 说明 |
|---|---|---|
| 自动构造 | --element-updates '{"subtitle":{"content":"新内容"}}' | 从 old spec + updates 自动生成精确的编辑 prompt |
| 手动提供 | --edit-prompt "在参考图基础上..." | Agent/用户直接写完整 prompt,跳过自动构造 |
两种方式可以同时用:--element-updates 更新 metadata 中的 spec,--edit-prompt 覆盖自动生成的 prompt。
---
四、回滚流程 (--rollback)
--rollback 3 --to-version 1 --session <ts>
│
▼
┌─ cmd_rollback_slide() ───────────────────────────────┐
│ │
│ 1. 加载 metadata.json │
│ 2. 找到 version=1 的 spec 和 image_snapshot │
│ 3. 有 elements? │
│ ├── YES → generate_prompt_from_spec(v1_spec) │
│ └── NO → 读 v1 的 prompt_file 原文 │
│ ├── 有 → 用原文 │
│ └── 无 → 用 v1 参考图构造视觉提示 │
│ 4. 以 v1 的图片作 reference (可选) │
│ 5. gpt-image-2 → 新图 │
│ 6. 备份当前图 → slide-03_v{old}.png │
│ 7. 新图 → slide-03.png │
│ 8. metadata.json 追加新 version: │
│ { action:"rollback", spec:<v1_spec>, │
│ reference_version:1 } │
│ 9. 重建 .pptx │
└───────────────────────────────────────────────────────┘注意:回滚不是直接恢复旧文件,而是用旧版本的 spec 重新生成一张新图,这样既保留了完整的版本历史,又确保了图片质量(旧图可能被压缩或损坏)。
---
五、摄取外部 PPTX (--ingest-pptx)
外部 deck.pptx
│
│ --ingest-pptx
▼
┌─ cmd_ingest_pptx() ──────────────────────────────────┐
│ │
│ 1. render_template.py → template_renders/<stem>/ │
│ 2. 创建 outputs/<ts>_<stem>/images/ │
│ 3. 复制 PNG: slide-NN_v0001.png + slide-NN.png │
│ 4. 写入 metadata.json (placeholder spec): │
│ { elements: {}, layout: "(待 Agent 分析填充)" } │
└───────────────────────────────────────────────────────┘
│
│ Agent 后续操作
▼
┌─ Agent ───────────────────────────────────────────────┐
│ Read 每页 slide-NN.png (多模态看懂内容) │
│ │
│ 对每页构建 slide_spec: │
│ { │
│ "elements": { │
│ "title": {"type":"heading", "content":"...", │
│ "position":"左上角", "style":"32pt", │
│ "color":"#ffffff"}, │
│ "body": {...} │
│ } │
│ } │
│ │
│ 更新 metadata.json → 之后走正常 --edit 流程 │
└───────────────────────────────────────────────────────┘---
六、核心数据结构
slide_spec
每页幻灯片的结构化描述,精确到每个可编辑元素:
{
"slide_number": 3,
"page_type": "content",
"layout": "两个卡片纵向排列",
"elements": {
"title": {
"type": "heading",
"content": "市场痛点",
"style": "48pt Bold 思源黑体",
"position": "左上角",
"color": "#00ff88"
},
"subtitle": {
"type": "subheading",
"content": "健康管理的两类割裂",
"style": "24pt Regular 思源黑体",
"position": "标题下方",
"color": "#ffffff"
},
"card_1": {
"type": "card",
"heading": "痛点一:高频无深度",
"body": "用户日均使用健康 App 3.2 次…",
"position": "中部偏左"
}
}
}元素字段说明:
| 字段 | 说明 | 示例 |
|---|---|---|
type | 元素语义类型 | heading, subheading, card, metric, decoration |
content | 元素内容文字 | "市场痛点" |
position | 页面位置 | "左上角", "标题下方", "中部偏左" |
style | 字体/样式描述 | "48pt Bold 思源黑体" |
color | 颜色 | "#00ff88" |
description | 装饰类元素的描述 | "深色渐变,左侧极光纹理" |
元素 ID 命名规范:
| 语义 | ID | 说明 |
|---|---|---|
| 主标题 | title | 每页唯一 |
| 副标题 | subtitle | 每页唯一 |
| 卡片 | card_1, card_2, … | 编号从 1 开始 |
| 数据指标 | metric_1, metric_2, … | 编号从 1 开始 |
| 背景装饰 | background | 装饰类,用 description 而非 content |
metadata.json
版本化的 slide_spec 存储:
{
"version": 1,
"title": "MediWise 商业计划书",
"style": "styles/dark-aurora.md",
"generated_at": "2024-05-23T14:30:52",
"slide_order": [1, 2, 3],
"slides": {
"3": {
"slide_number": 3,
"page_type": "content",
"current_version": 2,
"image_snapshot": "images/slide-03.png",
"versions": [
{
"version": 1,
"action": "generate",
"spec": { "elements": { "title": {...}, "subtitle": {...} } },
"prompt_file": "images/slide-03_v0001.txt",
"image_snapshot": "images/slide-03_v0001.png"
},
{
"version": 2,
"action": "edit",
"spec": { "elements": { "title": {...}, "subtitle": {"content": "医疗数据碎片化"} } },
"edit_instruction": "将副标题从'健康管理的两类割裂'改为'医疗数据碎片化'",
"prompt_file": "images/slide-03_v0002.txt",
"reference_version": 1,
"image_snapshot": "images/slide-03_v0002.png"
}
]
}
}
}版本链关系:
Slide 3:
v1 (generate) ──── spec₀ ──── slide-03_v0001.png
│
│ --edit: subtitle 被改了
▼
v2 (edit) ──── spec₁ ──── slide-03_v0002.png ← 当前
│
│ --rollback --to-version 1
▼
v3 (rollback) ─── spec₀ ──── slide-03_v0003.png ← 新当前images/slide-03.png 始终指向 current_version。
---
七、文件布局
outputs/<timestamp>/
├── images/
│ ├── slide-01.png # 当前(最新版本)
│ ├── slide-01_v0001.png # v1 快照
│ ├── slide-01_v0001.txt # v1 的 prompt
│ ├── slide-01_v0002.png # v2 快照
│ ├── slide-01_v0002.txt # v2 的 prompt
│ ├── slide-02.png
│ └── ...
├── metadata.json # 版本历史 + slide_spec
├── prompts.json # 兼容旧格式
└── <title>.pptx # 16:9 打包---
八、数据安全
- 原子写入:
metadata.json先写 temp 再 rename,不会因进程崩溃产生半截文件 - 版本快照:每次编辑前自动备份当前图为
_vXXXX.png,永不丢失 - 合并不覆盖:
--output指向已有 session 时合并新页,不覆盖已有 metadata - prompts.json 兼容:始终同步输出,保持与旧版脚本的兼容性
---
九、CLI 命令汇总
| 命令 | 作用 |
|---|---|
python3 scripts/generate_ppt.py --plan slides_plan.json --style styles/xx.md | 生成(内置风格) |
python3 scripts/generate_ppt.py --plan slides_plan.json --template-pptx xx.pptx --template-strict | 生成(模板克隆) |
python3 scripts/generate_ppt.py --plan slides_plan.json --style xx.md --slides 1,3 --output path/existing | 生成指定页,合并到已有 session |
python3 scripts/generate_ppt.py --edit 3 --session <ts> --element-updates '{"subtitle":{"content":"新"}}' | 修改第 3 页 |
python3 scripts/generate_ppt.py --edit 3 --session <ts> --edit-prompt "…" --element-updates '{…}' | 修改(手动 prompt + spec 更新) |
python3 scripts/generate_ppt.py --rollback 3 --to-version 1 --session <ts> | 回滚第 3 页到 v1 |
python3 scripts/generate_ppt.py --ingest-pptx path/to/deck.pptx | 摄取外部 PPTX |
python3 scripts/generate_ppt.py --list-sessions | 列出所有 session |
#!/bin/bash
##############################################################################
# gpt-image2-ppt-skills -- Claude Code / Codex Skill 安装脚本
#
# 把当前仓库内容拷贝到目标 skill 目录
# 并安装 Python 依赖、提示环境变量注入方式。
#
# 用法:bash install_as_skill.sh [--target auto|claude|codex|openclaw]
##############################################################################
set -e
RED='\033[0;31m'
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
BLUE='\033[0;34m'
NC='\033[0m'
print_info() { echo -e "${BLUE}(i) $1${NC}"; }
print_success() { echo -e "${GREEN}[OK] $1${NC}"; }
print_warning() { echo -e "${YELLOW}(!) $1${NC}"; }
print_error() { echo -e "${RED}[X] $1${NC}"; }
print_header() { echo ""; echo "========================================"; echo "$1"; echo "========================================"; echo ""; }
command_exists() { command -v "$1" >/dev/null 2>&1; }
TARGET="auto"
parse_args() {
while [[ $# -gt 0 ]]; do
case "$1" in
--target)
TARGET="${2:-}"
shift 2
;;
--target=*)
TARGET="${1#*=}"
shift
;;
*)
print_error "未知参数: $1"
echo "用法: bash install_as_skill.sh [--target auto|claude|codex|openclaw]"
exit 1
;;
esac
done
}
resolve_install_target() {
case "$TARGET" in
auto)
if [ -n "${CODEX_HOME:-}" ]; then
echo "codex"
elif [ -d "$HOME/.claude" ]; then
echo "claude"
elif [ -d "$HOME/.codex" ]; then
echo "codex"
else
echo "claude"
fi
;;
claude|codex|openclaw)
echo "$TARGET"
;;
*)
print_error "不支持的 target: $TARGET"
echo "可选值: auto | claude | codex | openclaw"
exit 1
;;
esac
}
resolve_skill_dir() {
case "$1" in
claude)
echo "$HOME/.claude/skills/gpt-image2-ppt-skills"
;;
codex)
echo "${CODEX_HOME:-$HOME/.codex}/skills/gpt-image2-ppt-skills"
;;
openclaw)
echo "$HOME/skills/gpt-image2-ppt"
;;
esac
}
resolve_agent_label() {
case "$1" in
claude)
echo "Claude Code"
;;
codex)
echo "Codex"
;;
openclaw)
echo "OpenClaw"
;;
esac
}
main() {
parse_args "$@"
print_header "gpt-image2-ppt-skills -- 安装"
INSTALL_TARGET="$(resolve_install_target)"
SKILL_DIR="$(resolve_skill_dir "$INSTALL_TARGET")"
AGENT_LABEL="$(resolve_agent_label "$INSTALL_TARGET")"
print_info "目标 agent: $AGENT_LABEL"
print_info "目标目录: $SKILL_DIR"
if [ -d "$SKILL_DIR" ]; then
print_warning "Skill 目录已存在: $SKILL_DIR"
read -p "是否覆盖?(y/N) " -n 1 -r
echo
if [[ ! $REPLY =~ ^[Yy]$ ]]; then
print_info "取消"
exit 0
fi
# 备份用户的 .env
if [ -f "$SKILL_DIR/.env" ]; then
cp "$SKILL_DIR/.env" "/tmp/gpt-image2-ppt.env.bak"
print_info "已备份现有 .env 到 /tmp/gpt-image2-ppt.env.bak"
fi
rm -rf "$SKILL_DIR"
fi
print_info "创建 Skill 目录..."
mkdir -p "$SKILL_DIR"
print_success "目录已创建"
print_info "复制项目文件..."
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# 拷贝核心文件,排除本地运行产物、凭据、缓存和测评代码。
rsync -a \
--exclude='.git' \
--exclude='outputs' \
--exclude='output' \
--exclude='outputs_cache' \
--exclude='results' \
--exclude='logs' \
--exclude='template_cache' \
--exclude='template_renders' \
--exclude='test_runs' \
--exclude='tmp' \
--exclude='venv' \
--exclude='.venv' \
--exclude='__pycache__' \
--exclude='.pytest_cache' \
--exclude='.env' \
--exclude='tests' \
--exclude='scripts/gen_demo_images.py' \
--exclude='scripts/gen_edit_case_images.py' \
"$SCRIPT_DIR/" "$SKILL_DIR/"
print_success "文件复制完成"
# 恢复备份的 .env
if [ -f "/tmp/gpt-image2-ppt.env.bak" ]; then
mv "/tmp/gpt-image2-ppt.env.bak" "$SKILL_DIR/.env"
print_success "已恢复用户 .env"
fi
print_info "检查 Python 环境..."
if ! command_exists python3; then
print_error "未找到 python3,请先安装 Python 3.8+"
exit 1
fi
print_success "Python: $(python3 --version)"
print_info "安装 Python 依赖..."
if command_exists pip3; then
pip3 install -q -r "$SKILL_DIR/requirements.txt"
else
pip install -q -r "$SKILL_DIR/requirements.txt"
fi
print_success "依赖安装完成"
print_header "环境变量配置提示"
if [ -f "$SKILL_DIR/.env" ]; then
print_info "已存在 skill 安装目录 .env(standalone CLI fallback),保留不改"
else
print_info "未自动创建 .env。推荐通过 agent 配置 / 系统环境变量注入 OPENAI_API_KEY"
print_info "standalone CLI 如需私有 env 文件,可复制 .env.example 后用 GPT_IMAGE2_PPT_ENV 指向它"
fi
print_header "安装完成"
print_success "已装到 $SKILL_DIR"
echo ""
print_info "下一步:"
print_info " 1. 如需 API 直连,通过 agent 配置 / 系统环境变量注入 OPENAI_API_KEY"
print_info " standalone CLI 可设置 GPT_IMAGE2_PPT_ENV=/path/to/private.env"
print_info " 2. 重启 $AGENT_LABEL 让 skill 生效"
if [ "$INSTALL_TARGET" = "codex" ]; then
print_info " 3. 在 Codex 里直接说:'帮我用 gpt-image2-ppt 生成一份 5 页 PPT'"
print_info " 如果当前 Codex 自带原生出图能力,可直接走原生路径,不必配置 OPENAI_API_KEY"
else
print_info " 3. 直接对 $AGENT_LABEL 说:'帮我用 gpt-image2-ppt 生成一份 5 页 PPT'"
fi
echo ""
print_info "冒烟测试(可选):"
print_info " cd $SKILL_DIR"
print_info " python3 scripts/generate_ppt.py --plan slides_plan.json --style styles/gradient-glass.md --slides 1"
echo ""
}
trap 'print_error "安装过程出错"; exit 1' ERR
main "$@"
requests>=2.31
python-dotenv>=1.0
python-pptx>=1.0
jsonschema>=4.0
pymupdf>=1.24
pillow>=10.0