
Promo Creator Skills
- 28 installs
- 74 repo stars
- Updated May 12, 2026
- kangarooking/promo-creator-skills
Create 60-90 second product demo videos from product information with storyboarding, asset planning and professional editing.
About
End-to-end workflow for producing product promo videos from brief through delivery. Coordinates storyboarding, asset sourcing, editing with HyperFrames and music synchronization.
- structured pipeline: brief -> storyboard -> assets -> editing -> BGM -> delivery
- sub-skill routing for each production stage with detailed specifications
Promo Creator Skills by the numbers
- 28 all-time installs (skills.sh)
- +1 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #980 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/kangarooking/promo-creator-skills --skill promo-creator-skillsAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 28 |
|---|---|
| repo stars | ★ 74 |
| Last updated | May 12, 2026 |
| Repository | kangarooking/promo-creator-skills ↗ |
What it does
Create 60-90 second product demo videos from product information with storyboarding, asset planning and professional editing.
Files
Promo Asset Producer — 两包素材生产
定位
你是宣传片素材制作人。根据确认后的分镜脚本,生产所有镜头需要的图片素材。
---
两包素材
Pack A — AI 生成 assets/pack-a/
使用当前环境可用的图片生成能力生成,例如 Codex imagegen、GPT Image、AutoGLM 图像生成或用户指定的图像模型。不要在没有对应工具时假装已经生成;如果图片生成不可用,输出可执行 prompt 并标注待生成。
可生成的素材类型
| 类型 | 适合 | Prompt 要点 |
|---|---|---|
| 产品界面模拟图 | Product Hero 镜头 | 深色/浅色编辑器、app 界面、dashboard |
| Logo 透明背景 | Hook / CTA 镜头 | 极简、矢量感、透明背景 |
| 概念插图 | Feature 镜头 | 抽象表达概念(速度、连接、安全) |
| 抽象视觉元素 | 转场 / 背景 | 光影、粒子、几何背景;避免无意义渐变球和 AI 装饰感 |
| 截图再设计 | 功能展示 | 把真实截图重构为统一视觉风格 |
| 数据可视化图 | Social Proof 镜头 | 柱状图、增长曲线、数据卡片 |
Prompt 工程规范
每个 prompt 必须包含 5 层信息:
1. [输出物] 生成一张横向 [具体是什么] 的图片。
2. [风格] [Apple 产品风格 / 瑞士极简 / 赛博科技 / 极简商务],[具体视觉特征]。
3. [构图] [比例] 横向构图,主体 [位置],[留白方向]。
4. [否定] 不要 [text/border/watermark/logo/页眉/页脚/装饰边框]。
5. [技术] 背景 [颜色/透明],输出 [比例]。各风格 Prompt 模板
Apple 发布会风:
生成一张横向产品界面截图。[产品名] 的 [功能/界面] 在深色背景上展示,风格像 Apple WWDC 演示截图。深色窗口、精致阴影、代码或 UI 清晰可读。
构图:16:9 横向,主体居中,上下留白 8%。
不要:真实可读文字(会模糊)、水印、Logo、边框。
背景:纯黑 #000,可以有微弱的蓝色/紫色辉光从产品边缘溢出。瑞士国际主义风:
生成一张横向信息图。主题是 [概念]。Swiss International Typographic Style,浅灰底 #fafaf8,单色强调色 [IKB蓝 #002FA7 / 柠檬黄 / 柠檬绿 / 安全橙],12 列网格对齐,直角模块,1px 发丝线,大量留白。
构图:[16:9 / 21:9] 横向,内容左对齐,右侧留白 30%。
不要:渐变、阴影、圆角、3D、霓虹、多色。赛博科技风:
生成一张横向科技界面图。[产品名] 的终端/CLI 界面在纯黑底上展示,等宽字体,终端绿 #00ff41 或品牌色文字,微弱扫描线纹理,Hacker 美学。
构图:16:9 横向,主体居中偏左。
不要:卡通、3D、明亮色彩、商务感。极简商务风:
生成一张横向产品展示图。[产品名] 的界面在纯白底上展示,干净、克制、商务感。细线分割,一个品牌色 [色值] 做微弱强调,大量留白。
构图:16:9 横向,主体居中。
不要:科技感装饰、渐变、阴影过重。Prompt 后缀(每个 prompt 尾部追加)
输出必须是 [16:9/21:9/1:1] [横/方]向构图,主体居中但保留边距,画面密度中等。只保留核心画面本身,不要生成页眉、页脚、标题、页码、角标、署名、装饰边框。多图同组追加:
这是一组图片中的一张,请保持与同组图片相同的画面比例、元素大小、边距、线条粗细和标注密度。图片比例规范
| 用途 | 推荐比例 |
|---|---|
| 全屏 hero / 产品主图 | 16:9 |
| 瑞士风顶部横幅 | 21:9 |
| 左右分屏主图 | 16:10 或 4:3 |
| 图文混排小图 | 3:2 或 3:4 |
| Logo / 图标 | 1:1(透明背景 PNG) |
命名规则
pack-a/
a-01-hook-logo.png ← Shot 1 的 Logo
a-02-problem-complex.png ← Shot 2 的概念图
a-03-product-ui.png ← Shot 3 的产品界面
a-04-feature-speed.png ← Shot 4 的功能插图---
Pack B — 网络搜索 assets/pack-b/
使用浏览器搜索下载真实素材。
可搜索的素材类型
| 类型 | 来源 | 注意 |
|---|---|---|
| 产品官方截图 | 官网 / GitHub | 注意版权 |
| 竞品界面截图 | 竞品官网 | 用于 Problem 段对比 |
| 开源照片 | Unsplash / Pexels | 确认 License |
| GitHub 界面 | github.com | Star 数、贡献者图 |
| 工具操作截图 | 实际工具页面 | 浏览器截图 |
下载规范
- 分辨率 ≥ 1920px 宽
- JPG 用于照片,PNG 用于界面/透明
- 记录每张图的来源 URL 和许可证
命名规则
pack-b/
b-02-premiere-ui.png ← Shot 2 的竞品截图
b-05-github-stars.png ← Shot 5 的 GitHub 截图
pack-b-sources.md ← 来源记录pack-b-sources.md 格式
## 素材来源记录
| 文件 | 来源 | 许可证 | 日期 |
|------|------|--------|------|
| b-02-premiere-ui.png | Adobe Premiere 官网截图 | Fair Use (对比用途) | 2026-05-11 |
| b-05-github-stars.png | github.com/xxx/xxx | 截图 | 2026-05-11 |---
素材计划表
输出 03-asset-plan.md,汇总所有素材需求:
## 素材计划
| # | Shot | 素材 | 来源 | 规格 | Prompt/来源 | 状态 |
|---|------|------|------|------|------------|------|
| 1 | Hook | Logo 透明背景 | Pack A | 1:1 PNG | [prompt] | ⏳ |
| 2 | Problem | 复杂编辑软件截图 | Pack B | 16:9 PNG | Adobe 官网 | 🔍 |
| 3 | Product | 产品界面模拟图 | Pack A | 16:9 PNG | [prompt] | ⏳ |
| 4 | Feature 1 | 功能概念图 | Pack A | 16:9 PNG | [prompt] | ⏳ |
| 5 | Social | GitHub Stars 截图 | Pack B | 16:9 PNG | GitHub | 🔍 |---
质量自检
P0 级(必须修)
- [ ] 图片比例畸形(非标准比例)
- [ ] AI 生图包含了文字/页眉/边框
- [ ] 风格与 brief 不一致
- [ ] 分辨率 < 1920px
P1 级(应该修)
- [ ] 同组图片风格不统一
- [ ] 留白不足,画面拥挤
- [ ] 背景色与选定风格不匹配
---
工作流
1. 读取确认后的 02-storyboard.md 2. 提取所有素材需求 3. Pack A: 按 prompt 规范逐个生成 AI 图片 4. Pack B: 使用浏览器搜索下载素材,记录来源 5. 输出 03-asset-plan.md + assets/pack-a/ + assets/pack-b/ 6. [PAUSE] 等用户确认素材质量
输出文件
outputs/promo-runs/<product>/
03-asset-plan.md
assets/
pack-a/ ← AI 生成图片
pack-b/ ← 网络搜索素材
pack-b-sources.md ← 来源记录.DS_Store
__pycache__/
*.pyc
.env
.env.*
*.log
# Local production runs are working artifacts.
# Curated public demos live in demos/.
outputs/
# Generated audio candidates can contain paid API output.
*.wav
*.flac
*.m4a
MIT License
Copyright (c) 2026 kangarooking
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
Promo Creator Skills
English · 中文
A production-grade skill pack for turning real products into credible launch films: product judgment, narrative structure, visual planning, HyperFrames editing, BGM direction, and delivery notes.
Most AI-generated promo videos fail for the same reason: they start with visuals before they understand the product. They decorate instead of deciding. They turn a repo, a SaaS page, or a new feature into generic motion graphics.
Promo Creator Skills takes the opposite route. It treats a promo film as a sequence of production decisions: what the product is, why it matters, what the viewer must understand first, which real assets prove it, how each shot moves, and what kind of music actually supports the cut. The result is not a magic button. It is a disciplined agent workflow for making product videos that feel specific, inspectable, and worth shipping.
Use it when the goal is not "make something flashy", but "make the product legible, credible, and worth a second look in 60-90 seconds."
npx skills add kangarooking/promo-creator-skillsIf your skill runner does not support npx skills add, copy the promo-* folders into your agent skills directory.
---
Demo Gallery
WorkBuddy OPC · Commercial Product Launch
Modern product-launch BGM direction: beat, bass, hook, and CTA momentum. Useful when the video should feel like a launch film rather than background UI music.
WorkBuddy OPC · Clean UI Demo Groove
Software-demo BGM direction: cleaner rhythm, UI-card timing, workflow motion, and enough energy to carry real interface shots.
Cangjie Skill · Open Source Project Promo
A 90-second Swiss International-style promo for an open-source meta-skill that compiles books into callable Agent Skills. The point is not louder claims; it is making a valuable method visible: book knowledge becomes triggers, steps, boundaries, and tests an agent can actually use.
GitHub README sanitizes external player iframes, so Bilibili players cannot be embedded directly here. Use the links above for real-time playback on Bilibili.
---
What It Does
This pack is built around one principle: a good product video is a compressed product argument. The visuals are there to carry that argument, not to hide the absence of one.
| Skill | Job | Output |
|---|---|---|
promo-brief | Research the product and choose the narrative direction | 01-brief.md |
promo-storyboard | Write shot-by-shot visual and motion specs | 02-storyboard.md |
promo-asset-producer | Plan and gather generated assets plus real product material | 03-asset-plan.md, assets/ |
promo-editor | Build the HyperFrames edit and render the MP4 | 04-edl.md, master-edit.html, final/promo.mp4 |
promo-music-maker | Design BGM prompts, hit points, and product-promo music lanes | 06-music-plan.md, assets/bgm/ |
promo-workflow | Orchestrate the whole staged pipeline | delivery-ready run folder |
The workflow is deliberately staged. A good promo is not a one-shot prompt; it is a chain of decisions:
Product / repo / URL
-> brief
-> storyboard
-> asset plan
-> edit decision list
-> HyperFrames master edit
-> BGM direction and candidates
-> final MP4
-> delivery notesEach stage leaves an artifact behind, so you can challenge the reasoning before paying the cost of rendering, regenerating assets, or burning music credits.
---
Why This Is Different
- It starts with product strategy: the brief decides audience, positioning, narrative shape, and proof points before any frame is designed.
- It designs shots, not vibes: the storyboard specifies frame composition, typography, timing, motion, and asset needs.
- It prefers real product evidence: screenshots, docs, GitHub data, and interface details carry more weight than decorative AI imagery.
- It treats BGM as part of editing: music prompts include genre, rhythm, bass, hook, energy curve, and hit points instead of vague "premium tech" language.
- It keeps the work inspectable: every phase produces Markdown or source files that can be reviewed, corrected, and reused.
---
BGM That Does Not Collapse Into Generic AI Ambience
The music skill includes a hard-won product-promo rule: visual minimalism is not the same as ambient music.
For software product videos, the default pair is now:
- Commercial Product Launch: 118 BPM, tight electronic drums, sidechained synth bass, bright plucked synth hook.
- Clean UI Demo Groove: 112 BPM, syncopated digital percussion, rubbery synth bass, short glass-pluck motif.
This avoids the common failure mode where every option becomes minimal / ambient / soft pulse / glass ticks and sounds interchangeable.
Mureka generation is supported through scripts/mureka.py:
export MUREKA_API_KEY="..."
export MUREKA_API_URL="https://api.mureka.cn" # optional, for the China endpoint
python scripts/mureka.py instrumental \
--prompt "<validated product-promo BGM prompt>" \
-n 2 \
--format mp3 \
--output outputs/promo-runs/<run-id>/assets/bgm/---
Usage
Ask your agent:
Use promo-workflow to make a 60-second Apple-style product promo for this GitHub repo:
https://github.com/owner/projectOr:
Use the promo skills to create a launch video for my SaaS. Pause after the brief and storyboard.For existing footage or screenshots:
Use promo-editor and promo-music-maker. I already have the storyboard and assets.---
Repository Structure
promo-creator-skills/
├── README.md # Chinese default README
├── README.en.md # English README
├── SKILL.md # skill pack entrypoint
├── promo-brief/ # product research and creative brief
├── promo-storyboard/ # shot-by-shot visual specs
├── promo-asset-producer/ # generated and sourced asset planning
├── promo-editor/ # HyperFrames edit and render workflow
├── promo-music-maker/ # product-promo BGM prompt and Mureka workflow
├── promo-workflow/ # orchestration skill
├── references/
│ ├── mureka_prompt_guide.md
│ └── product_promo_bgm_prompting.md
└── scripts/
└── mureka.pySKILL.md at the repository root is the full orchestration entrypoint: it can run the main workflow on its own, while the promo-* folders keep deeper stage-specific guidance for focused execution.
---
Requirements
- Node.js 22+
- HyperFrames CLI for HTML-to-video work
- FFmpeg available on
PATH - Python 3.10+
requestsfor Mureka API calls- Optional:
MUREKA_API_KEYfor BGM generation
The skills can still design BGM prompts without an API key; they only call Mureka when explicitly asked to generate audio.
---
Design Principles
- Intermediate files over black boxes: every major decision lands in Markdown before render.
- Real product signal first: UI screenshots, docs, GitHub data, and brand assets beat decorative filler.
- Storyboard before assets: image generation without a shot plan is expensive randomness.
- Music has a job: product videos need rhythm, hook, bass movement, and section contrast.
- Pause points are part of quality: good agent workflows invite correction before the expensive step.
---
Limitations
- This is a skill pack, not a hosted video editor.
- Final visual quality depends on available product assets and the agent's render environment.
- Mureka output can vary; generate 2-3 candidates and pick by watching the video, not by listening in isolation.
- Large local production runs are ignored by git; curated public examples should be uploaded to a video platform and linked from the README.
---
License
MIT.
Promo Creator Skills
English · 中文
一套面向真实产品宣传片的 production-grade Skills:产品判断、叙事结构、视觉规划、HyperFrames 剪辑、BGM 设计和交付清单,全链路跑通。
大多数 AI 宣传片失败,不是因为动效不够多,而是因为一开始就跳过了产品判断。它们先做画面,再临时找理由;先堆光效,再假装有叙事。最后出来的东西看起来“像视频”,但不像一个真正懂产品的人会拿出去发布的片子。
Promo Creator Skills 反过来做。它把一支宣传片拆成一组严肃的生产决策:这个产品到底是什么,为什么值得看,观众第一眼应该理解什么,哪些真实素材能证明它,每个镜头如何推进信息,音乐如何托住剪辑节奏。它不是魔法按钮,而是一套让 agent 按专业制作流程工作的技能系统。
适合的目标不是“做得炫一点”,而是:在 60-90 秒内,把产品讲清楚、讲可信,并让人愿意继续了解。
npx skills add kangarooking/promo-creator-skills如果你的 skill runner 不支持 npx skills add,可以手动把 promo-* 文件夹复制到你的 agent skills 目录。
---
案例展示
WorkBuddy OPC · Commercial Product Launch
更像现代产品发布广告片:有 beat、bass、hook 和 CTA 推进感。适合产品上架、新能力发布、商业展示。
WorkBuddy OPC · Clean UI Demo Groove
更像高级软件 UI 演示片:节奏干净,服务界面、卡片矩阵、流程线和时间线卡点。
Cangjie Skill · 开源项目宣传片
一支 90 秒瑞士国际主义风开源项目宣传片。Cangjie Skill 的价值不在于“又一个读书总结工具”,而在于把高价值书籍的方法论编译成 Agent 可以调用的触发条件、执行步骤、边界和测试。好的开源项目不是把 README 录一遍,而是把它真正值得被记住的系统讲清楚。
GitHub README 会清洗外部播放器 iframe,不能直接内嵌 B 站播放器;点击链接可在 Bilibili 实时播放。
---
它能做什么
这套 Skills 的基本判断是:好的产品宣传片,本质是一段被压缩过的产品论证。 画面、动效、音乐都应该服务这个论证,而不是掩盖产品表达的空心。
| Skill | 职责 | 产物 |
|---|---|---|
promo-brief | 产品研究、卖点提炼、叙事定位 | 01-brief.md |
promo-storyboard | 逐镜头分镜、画面、文案、动效说明 | 02-storyboard.md |
promo-asset-producer | 生成素材与真实产品素材规划 | 03-asset-plan.md, assets/ |
promo-editor | HyperFrames 剪辑、动效、渲染 | 04-edl.md, master-edit.html, final/promo.mp4 |
promo-music-maker | BGM 方向、卡点表、Mureka Prompt | 06-music-plan.md, assets/bgm/ |
promo-workflow | 串联完整制作流程 | 可交付 run 目录 |
整套流程不是一发入魂,而是把宣传片拆成可检查、可修改、可复用的阶段:
产品 / 仓库 / URL
-> 创意简报
-> 逐镜头分镜
-> 素材计划
-> 编辑决策表
-> HyperFrames 主剪辑
-> BGM 方向和候选
-> 最终 MP4
-> 交付清单每个阶段都会留下文件,所以你可以在真正昂贵的步骤之前纠偏:渲染之前改分镜,生图之前改素材计划,烧音乐额度之前改 BGM 方向。
---
为什么它不只是 Prompt 合集
- 先做产品判断:brief 阶段先确定受众、定位、叙事结构和证明点,再进入视觉。
- 写的是镜头,不是氛围词:storyboard 会拆到构图、字号、时间、动效、素材需求。
- 真实产品信号优先:截图、文档、GitHub 数据、界面细节,比 AI 装饰图更有价值。
- BGM 属于剪辑,不是背景噪音:音乐 prompt 明确 genre、rhythm、bass、hook、energy curve 和 hit points。
- 过程可检查:每一步都有 Markdown 或源文件,方便复盘、修改和复用。
---
BGM:不再生成那种“泛科技垫底音乐”
这次专门把 BGM 设计升级了:Apple 风画面不等于 ambient BGM。产品宣传片需要节奏、hook、bass movement 和段落变化。
软件产品宣传片默认先走两条已验证方向:
- Commercial Product Launch:118 BPM,电子鼓、sidechain bass、明亮 pluck hook,更像产品发布广告片。
- Clean UI Demo Groove:112 BPM,syncopated digital percussion、rubbery synth bass、短促 glass-pluck motif,更像高级 UI 演示片。
这样不会再出现多个候选都像 minimal / ambient / soft pulse / glass ticks 的问题。
Mureka 生成脚本:
export MUREKA_API_KEY="..."
export MUREKA_API_URL="https://api.mureka.cn" # 国内站可选
python scripts/mureka.py instrumental \
--prompt "<validated product-promo BGM prompt>" \
-n 2 \
--format mp3 \
--output outputs/promo-runs/<run-id>/assets/bgm/没有 API key 也可以使用这个 skill 设计 BGM prompt;只有用户明确要求实际生成音乐时才调用 Mureka。
---
使用方式
对 agent 说:
用 promo-workflow 给这个 GitHub 仓库做一支 60 秒 Apple 风产品宣传片:
https://github.com/owner/project或者:
用这套 promo skills 给我的 SaaS 做发布视频。先出 brief 和 storyboard,我确认后再继续。如果你已经有素材:
直接用 promo-editor 和 promo-music-maker。我已经有分镜和截图。---
仓库结构
promo-creator-skills/
├── README.md # 中文默认介绍
├── README.en.md # English README
├── SKILL.md # skill pack 入口
├── promo-brief/ # 产品研究与创意简报
├── promo-storyboard/ # 逐镜头分镜
├── promo-asset-producer/ # 素材生产与素材计划
├── promo-editor/ # HyperFrames 剪辑与渲染
├── promo-music-maker/ # BGM Prompt 与 Mureka 工作流
├── promo-workflow/ # 总控流程
├── references/
│ ├── mureka_prompt_guide.md
│ └── product_promo_bgm_prompting.md
└── scripts/
└── mureka.py仓库根目录的 SKILL.md 是完整总控入口:只读它也能跑通主流程;各个 promo-* 子 skill 保留更细的阶段规范,方便单独调用和深度执行。
---
环境依赖
- Node.js 22+
- HyperFrames CLI
- FFmpeg
- Python 3.10+
requestsPython 包- 可选:
MUREKA_API_KEY,用于实际生成 BGM
---
设计原则
- 中间文件优先:每个关键决策先落 Markdown,再进入渲染。
- 真实产品信号优先:UI 截图、文档、GitHub 数据、品牌素材,优先级高于装饰图。
- 先分镜,后素材:没有 shot plan 的生图只是烧钱随机数。
- 音乐必须服务画面:产品片需要 rhythm、hook、bass 和 section contrast。
- 暂停点就是质量控制:在贵的步骤之前,让用户有机会纠偏。
---
边界
- 这是 skill pack,不是托管视频编辑器。
- 最终质量取决于产品素材质量和本地渲染环境。
- Mureka 生成有随机性,建议一次生成 2-3 条,放回视频里看,而不是单独听音频决定。
- 本地完整生产过程放在
outputs/,默认不进 git;公开案例建议上传到视频平台后在 README 引用。
---
License
MIT.
Music Prompt Crafting Guide
Comprehensive guide to writing effective music prompts for Mureka AI.
The Golden Rule: DESCRIBE, Don't Command
❌ "Create an energetic pop song with drums"
✅ "energetic pop, driving four-on-the-floor drums, bright synth hooks, 128 BPM, female vocal, festival anthem vibe"The AI responds to descriptions of the music, not instructions to "make" or "create" it.
---
Design a Dynamic Arc, Not a Static Description
The most common reason AI-generated songs sound flat is that the prompt describes one mood level throughout. Great songs are tension/release journeys — design yours explicitly.
Express the arc as a mood progression string in your prompt:
✅ "sparse and intimate opening → rising tension → full cathartic chorus → stripped-back bridge → bigger final chorus"
✅ "melancholic and sparse → building urgency → explosive release → quiet resolution"The AI interprets this as an emotional journey across the song. Without it, every section gets the same energy and density.
---
Effective Prompt Structure
Minimum Viable Prompt (always include these)
[genre + sub-genre], [mood/emotion], [tempo/BPM], [key instruments], [vocal style]Standard Prompt (recommended for good results)
Genre: [specific genre + era, e.g., "90s trip-hop"]
Mood: [2-3 descriptors, e.g., "melancholic, introspective, nocturnal"]
Tempo: [BPM or description, e.g., "85 BPM, slow groove"]
Instruments: [3-5 key instruments, e.g., "turntable scratches, Rhodes piano, upright bass"]
Vocals: [style, e.g., "breathy female vocals, intimate delivery"]
Scene: [usage context, e.g., "late-night city driving"]Production Brief (for maximum control)
Genre: trip-hop | Era: mid-90s Bristol
BPM: 85 | Key: D minor
Mood: melancholic → building tension → cathartic release
Lead: Rhodes piano, tremolo
Rhythm: breakbeat, vinyl crackle texture
Bass: deep sub-bass, Moog-style
Texture: tape saturation, lo-fi warmth
Vocals: breathy female, close-mic intimacy
Structure: Intro(8bars) / Verse / Chorus / Verse / Bridge / Chorus / Outro
Avoid: auto-tune, bright synths, four-on-the-floor kick
Reference: Portishead "Roads" (breakbeat texture, Rhodes tone)Note on references: Only use specific songs as sonic benchmarks when all other parameters are already specified. Different from "sound like [artist]" — use for named production qualities (e.g., "Roads-style breakbeat texture").
---
What Makes Prompts FAIL (Top 7 Mistakes)
| # | Mistake | Why It Fails | Fix |
|---|---|---|---|
| 1 | Vague prompts ("nice pop song") | AI defaults to the statistical average — generic, forgettable | Be ruthlessly specific: sub-genre + era + mood + instruments + BPM |
| 2 | Contradictions ("slow and relaxing, high energy, 160 BPM") | Conflicting signals make the AI unpredictable | Check every descriptor agrees with the mood. Pick one direction |
| 3 | "Sound like [famous artist]" | Copyright risk + AI interprets literally, often misses the point | Describe the qualities you like: "warm analog synths, driving bass, 80s production style" |
| 4 | Too many words per lyric line | AI rushes through words → slurred, unnatural vocals | Keep lines ≤10 words. Short lines = better vocal delivery |
| 5 | No structure tags in lyrics | Song has no shape — verse/chorus blur together | Always use [Verse], [Chorus], [Bridge], [Outro] tags |
| 6 | Rewriting entire prompt between iterations | Can never isolate what improved (or worsened) the output | Change ONE element at a time. A/B test systematically |
| 7 | Ignoring negative prompts | Unwanted elements creep in (auto-tune, trap hi-hats, reverb) | Explicitly state what to avoid: "no auto-tune, avoid heavy reverb" |
---
Effective vs Ineffective — Side-by-Side Examples
Example 1: Pop Song
❌ "A pop song about love that sounds good"
✅ "bright synth-pop, uplifting, 120 BPM, arpeggiated synths, punchy electronic drums, female vocal with light reverb, 2020s clean production, summer anthem feel"Example 2: Lo-Fi Background
❌ "lofi music for studying"
✅ "lo-fi hip-hop, warm and mellow, 75 BPM, dusty vinyl crackle, jazzy Rhodes chords, muted boom-bap drums, no vocals, late-night study session atmosphere"Example 3: Cinematic
❌ "epic movie music"
✅ "cinematic orchestral, tension building to triumphant climax, 95 BPM, strings staccato → legato swell, French horns, timpani rolls, choir in final section, Hans Zimmer-style layered percussion"Example 4: Rock Song
❌ "energetic rock song, male vocals, guitar solo"
✅ "alternative rock, energetic and raw, 140 BPM, distorted electric guitar riffs, driving bass line, punchy drums, raspy male vocals, anthemic chorus, guitar solo section, garage rock aesthetic, festival anthem energy"Example 5: Traditional Chinese
❌ "Chinese music, sad"
✅ "Chinese traditional guofeng, melancholic and nostalgic, 60 BPM, dizi bamboo flute lead melody, guzheng plucked strings, subtle erhu, misty atmosphere, Jiangnan water town imagery, rain and mist soundscape, instrumental only"---
Lyrics Writing: What Separates Good from Bad
Structure Tags (always use these)
[Intro]
[Verse]
[Pre-Chorus]
[Chorus]
[Bridge]
[Break]
[Outro]Standard structure order: Verse → Chorus → Verse → Chorus → Bridge → Final Chorus → Outro. The Bridge always appears after the second chorus — never before the first.
If your generated bridge sounds like a second verse, regenerate it with:
python generate_lyrics.py extend "<existing lyrics>" "write a contrasting bridge that shifts perspective, strips back to a single instrument, and sets up the final chorus"Golden Rules for Lyrics That Sing Well
1. Keep lines short — 6-10 words per line. Long lines get rushed.
❌ "I've been walking through the streets of this old town thinking about everything we used to do together"
✅ "Walking through the old town streets\n Thinking of what we used to be"2. Match syllable count across verse lines — Creates natural rhythm.
✅ "Shadows fall on empty streets" (7 syllables)
"Whispers lost in evening heat" (7 syllables)
"Dancing lights through window panes" (7 syllables)3. Use rhyme patterns intentionally — ABAB or AABB, not random.
✅ [Verse]
The city sleeps beneath the stars (A)
While dreamers chase the fading light (B)
We trace our names on passing cars (A)
And disappear into the night (B)4. Chorus should be simpler and more repetitive than verses — Fewer words, not more. Whitespace and repetition create impact; repetition IS the melody.
5. Don't over-explain in lyrics — Imagery > exposition.
---
Writing a Memorable Hook
The hook is the most important line in your song. Get it right:
1. Length and singability
4-8 words, singable on first listen, usually contains the song's title.
✅ "I will always love you" — simple, universal, title, singable
✅ "Rolling in the deep" — 4 words, vivid, singable
❌ "I feel the way I feel when I think about our story" — too long, too vague2. End chorus lines on open vowels
AI vocals hold the last syllable of each line. Open vowels (oh, ah, ay, ee) sustain beautifully. Closed consonants (mm, th, ff, ss) sound awkward when held.
✅ "Let me go" → ends on "oh" — sustains well
❌ "Let me breathe" → ends on closed "th" sound — awkward to hold3. Chorus density
A chorus should have FEWER words than a verse, not more. The space around the hook gives it impact.
❌ "I feel sad because you left me and now I'm alone"
✅ "Empty chair across the table\n Coffee cold, the morning grey"---
Auto-Generate Lyrics First, Then Refine
# Generate lyrics from a concept
python generate_lyrics.py generate "a bittersweet farewell song, two old friends parting ways after summer"
# Extend if you need more sections
python generate_lyrics.py extend "[Verse]\nThe last light paints the pier in gold..."---
Iteration Strategy (How Pros Refine)
1. First generation: Use your best-guess prompt + n=3 2. Listen to all choices: Note what's good and what's off 3. Adjust ONE element: If rhythm is wrong → change BPM/drums description. If mood is off → change mood descriptors 4. Re-generate: Same lyrics, tweaked prompt 5. Compare: Does the change improve or worsen? 6. Repeat until satisfied
What to Listen For
For vocal songs:
| Symptom | Fix |
|---|---|
| Vocals feel rushed / words swallowed | Shorten lyric lines, reduce syllables per line |
| No energy build between verse and chorus | Add mood progression arc to prompt (e.g., "sparse → full cathartic release") |
| Hook doesn't stick | Simplify chorus to 4-8 words, repeat title phrase, check lines end on open vowels |
| Tempo feels wrong | Adjust BPM ±10 and regenerate |
| Listed instruments not audible | Verify each instrument is named explicitly in the prompt |
For instrumentals / ambient:
| Symptom | Fix |
|---|---|
| Tempo feels wrong | Adjust BPM ±10 |
| Mood doesn't match intent | Audit all mood descriptors for internal consistency |
| Instruments missing | Verify each is named explicitly in the prompt |
Never rewrite the entire prompt at once. You'll lose track of what works.
---
Production Checklist (Before You Generate)
Before hitting generate, verify:
- [ ] Genre is specific: Not just "pop" but "synth-pop, 2020s, clean production"
- [ ] Mood is consistent: No contradictions (slow + energetic = confused AI)
- [ ] BPM is set: Even approximate ("~90 BPM, slow groove") helps
- [ ] 3-5 instruments listed: Gives the AI sonic anchors
- [ ] Vocal style specified: Or "no vocals" / "instrumental only"
- [ ] Lyrics have structure tags: [Verse], [Chorus], [Bridge], [Outro]
- [ ] Lines are short: ≤10 words per line
- [ ] Avoid list included: What you DON'T want (auto-tune, trap hi-hats, etc.)
- [ ] N > 1: Generate 2-3 choices and pick the best. Never rely on a single generation
Product Promo BGM Prompting
Use this reference when a product promo BGM feels too generic, too ambient, too "AI tech background", or too similar across options.
Core Diagnosis
Visual style and music genre are not the same thing.
Apple-style visuals often means black stage, clean typography, and product focus. It does not always mean ambient pads, glass ticks, no drums, and weak energy. A product launch or product demo BGM usually needs one or more of:
- a clear rhythm bed
- a memorable but non-vocal hook
- bass movement
- section contrast
- commercial polish
- a satisfying final lift
Anti-Convergence Rule
Do not create alternatives by only changing BPM, a few adjectives, or the same instrument family.
Each option must differ on at least 4 of these 6 axes:
| Axis | Examples |
|---|---|
| Genre | electro-pop product launch, indie electronic, corporate cinematic, future garage, synthwave, minimal house |
| Rhythm | four-on-the-floor, half-time groove, syncopated clicks, breakbeat, no percussion |
| Bass | warm sub pulse, sidechained synth bass, plucked bass, cinematic low strings |
| Hook | synth pluck motif, piano motif, marimba/glass motif, string ostinato, bass riff |
| Energy curve | slow build, early hook, mid-film lift, final drop-out, constant drive |
| Instrument palette | drums + bass + synth hook, piano + strings, modular synth + percussion, guitar + electronic drums |
If two prompts share more than two of genre, rhythm, bass, hook, and instrument palette, they are not valid alternatives.
Better Product Promo Lanes
Validated Default Pair
For software product promo videos similar to the WorkBuddy OPC film, start with these two lanes first:
1. Commercial Product Launch: use when the video should feel like a modern product release ad. This was validated by promo-v2-commercial-launch-01.mp4. 2. Clean UI Demo Groove: use when the video should closely support UI screens, cards, workflows, and product operation. This was validated by promo-v2-ui-groove-02.mp4.
This pair is a better default than ambient tech music because it creates two genuinely different but relevant directions:
| Direction | Main job | Sonic identity |
|---|---|---|
| Commercial Product Launch | Product ad energy, launch polish, CTA momentum | 118 BPM, tight electronic drums, sidechained synth bass, plucked synth hook |
| Clean UI Demo Groove | UI timing, cards, workflow motion, clean software feel | 112 BPM, syncopated digital percussion, rubbery synth bass, short glass-pluck motif |
When generating first-round BGM options, create this validated pair before trying quieter ambient, cinematic, or emotional lanes.
Lane 1: Commercial Product Launch
Use when the user wants something closer to modern product ads, SaaS launch videos, or product showcase reels.
modern commercial product launch music, 118 BPM, instrumental only, polished and confident, tight electronic drums, sidechained synth bass, bright plucked synth hook, warm pads, clean impact hits. Structure: sparse branded intro for 0-6s, groove enters at 6s, short impact at 14s, full product-demo beat from 15-34s, stronger lift from 34-54s, clean final resolve at 54-60s. Designed for software UI, card animations, and product reveal cuts. No vocals, no lyrics, no ambient-only bed, no sleepy pads, no trap, no cinematic trailer booms.Lane 2: Clean UI Groove
Use when the video has UI screens, dashboards, cards, and workflow animations that need momentum.
clean UI demo groove, 112 BPM, instrumental only, precise and upbeat but not loud, muted kick, crisp snare snaps, syncopated digital percussion, rubbery synth bass, short glass-pluck motif, light airy pads. Structure: minimal boot-up intro, groove starts at 6s, UI reveal hit at 15s, card grid rhythm from 24s, tighter workflow pulse from 34s, wider final lift from 44s, quick clean ending. No vocals, no lyrics, no slow ambient wash, no generic corporate piano, no EDM festival drop.Lane 3: Founder Film With Momentum
Use when the story starts with pressure or loneliness but still needs product-commercial energy.
emotional founder-product film score with modern electronic groove, 98 BPM, instrumental only, focused and hopeful, soft piano motif, warm synth bass, brushed electronic drums, subtle strings, glass accents. Structure: lonely sparse intro, tension under founder overload at 6s, product reveal opens harmony at 15s, rhythm becomes organized at 24s, stronger teamwork movement at 34s, warm confident lift at 44s, resolved final phrase at 54s. No vocals, no lyrics, no sad piano ballad, no ambient-only pad, no heavy trailer drums.Lane 4: Premium Tech Pop Instrumental
Use when the visuals are clean but the user wants more of a finished ad-track feel.
premium tech-pop instrumental for a product showcase, 120 BPM, instrumental only, sleek, optimistic, commercially polished, punchy electronic drums, smooth synth bass, catchy plucked synth motif, soft chord stabs, subtle risers. Structure: 0-6s branded intro with filtered hook, 6-15s beat builds, 15-24s product UI groove, 24-34s hook variation for expert cards, 34-44s bigger rhythmic lift for parallel workflow, 44-54s open uplifting section, 54-60s short confident outro. No vocals, no lyrics, no sleepy ambient, no ukulele, no stock corporate jingle, no aggressive EDM.Lane 5: Executive Cinematic With Beat
Use when the brand needs credibility and polish, but not a flat corporate bed.
executive cinematic product score with restrained beat, 102 BPM, instrumental only, premium, serious, forward-moving, low strings, soft piano pulses, tight hybrid percussion, controlled synth bass, clean metallic hits. Structure: serious opening, tension at 6s, brand hit at 14s, credible UI reveal at 15s, measured card rhythm at 24s, stronger process momentum at 34s, broader confident lift at 44s, resolved final cadence at 54s. No vocals, no lyrics, no cheesy corporate piano, no ambient-only bed, no trailer booms, no trap hats.Prompt Rules For Mureka
- Stay under 1024 characters for generation.
- Put genre, BPM, rhythm, bass, and hook in the first sentence.
- Include "instrumental only" and "no vocals, no lyrics".
- Avoid overusing
minimal,ambient,pads,glass,ticks,Apple,premium techin every option. - For product demos, do not ban drums by default. Ban only the wrong drums:
no trap hats,no EDM festival drop,no trailer booms. - Use time-coded structure only after the sonic identity is clear.
"""Mureka AI music generation CLI — single entry point for all operations."""
import argparse
import json
import os
import sys
import time
import requests
# ---------------------------------------------------------------------------
# API Client
# ---------------------------------------------------------------------------
API_BASE = os.getenv("MUREKA_API_URL", "https://api.mureka.ai").rstrip("/")
POLL_INTERVAL = 5
POLL_TIMEOUT = 600
def get_api_key():
key = os.getenv("MUREKA_API_KEY")
if not key:
print("Error: MUREKA_API_KEY is not set", file=sys.stderr)
sys.exit(1)
return key
def headers(api_key=None):
key = api_key or get_api_key()
return {"Authorization": f"Bearer {key}", "Content-Type": "application/json"}
def post_json(path, payload, api_key=None, timeout=60):
url = f"{API_BASE}{path}"
resp = requests.post(url, json=payload, headers=headers(api_key), timeout=timeout)
if not resp.ok:
raise requests.HTTPError(f"{resp.status_code} {resp.reason}: {resp.text}", response=resp)
return resp.json()
def get_json(path, api_key=None, timeout=30):
url = f"{API_BASE}{path}"
resp = requests.get(url, headers=headers(api_key), timeout=timeout)
if not resp.ok:
raise requests.HTTPError(f"{resp.status_code} {resp.reason}: {resp.text}", response=resp)
return resp.json()
def upload_file_api(file_path, purpose, api_key=None):
key = api_key or get_api_key()
url = f"{API_BASE}/v1/files/upload"
with open(file_path, "rb") as f:
resp = requests.post(
url,
headers={"Authorization": f"Bearer {key}"},
files={"file": f},
data={"purpose": purpose},
timeout=120,
)
resp.raise_for_status()
return resp.json()
def poll_task(query_path, task_id, api_key=None, interval=POLL_INTERVAL, timeout=POLL_TIMEOUT):
key = api_key or get_api_key()
path = f"{query_path}/{task_id}"
deadline = time.time() + timeout
terminal = {"succeeded", "failed", "timeouted", "cancelled"}
while True:
data = get_json(path, api_key=key)
status = data.get("status", "")
print(f" [{status}] task {task_id}", file=sys.stderr)
if status in terminal:
if status != "succeeded":
reason = data.get("failed_reason", status)
raise RuntimeError(f"Task {task_id} ended with status: {status}. Reason: {reason}")
return data
if time.time() > deadline:
raise RuntimeError(f"Task {task_id} timed out after {timeout}s (last status: {status})")
time.sleep(interval)
def download_audio(url, output_path):
resp = requests.get(url, timeout=120)
resp.raise_for_status()
os.makedirs(os.path.dirname(os.path.abspath(output_path)), exist_ok=True)
with open(output_path, "wb") as f:
f.write(resp.content)
return output_path
def download_choices(result, args):
"""Download audio files into output directory."""
choices = result.get("choices", [])
if not choices:
print("No results generated.", file=sys.stderr)
sys.exit(1)
out_dir = args.output
os.makedirs(out_dir, exist_ok=True)
format_key = {"mp3": "url", "flac": "flac_url", "wav": "wav_url"}[args.format]
for choice in choices:
idx = choice.get("index", 0)
url = choice.get(format_key) or choice.get("url")
duration_ms = choice.get("duration", 0)
if len(choices) > 1:
filename = f"audio_{idx}.{args.format}"
else:
filename = f"audio.{args.format}"
output_path = os.path.join(out_dir, filename)
print(f"Downloading choice {idx} ({duration_ms / 1000:.1f}s) → {output_path}", file=sys.stderr)
download_audio(url, output_path)
print(output_path)
# ---------------------------------------------------------------------------
# Subcommands
# ---------------------------------------------------------------------------
def cmd_song(args):
"""Generate a song with lyrics and vocals."""
payload = {"lyrics": args.lyrics, "model": args.model}
if args.prompt:
payload["prompt"] = args.prompt
if args.reference_id:
payload["reference_id"] = args.reference_id
if args.vocal_id:
payload["vocal_id"] = args.vocal_id
if args.melody_id:
payload["melody_id"] = args.melody_id
if args.n:
payload["n"] = args.n
# Save lyrics and prompt to output directory
os.makedirs(args.output, exist_ok=True)
lyrics_path = os.path.join(args.output, "lyrics.txt")
with open(lyrics_path, "w", encoding="utf-8") as f:
f.write(args.lyrics)
if args.prompt:
f.write(f"\n\n---\nPrompt: {args.prompt}\n")
print(f"Lyrics saved → {lyrics_path}", file=sys.stderr)
print("Submitting song generation task...", file=sys.stderr)
task = post_json("/v1/song/generate", payload)
task_id = task["id"]
print(f"Task ID: {task_id}", file=sys.stderr)
print("Polling for completion...", file=sys.stderr)
result = poll_task("/v1/song/query", task_id,
interval=args.poll_interval, timeout=args.poll_timeout)
download_choices(result, args)
def cmd_instrumental(args):
"""Generate an instrumental track."""
payload = {"model": args.model}
if args.prompt:
payload["prompt"] = args.prompt
if args.instrumental_id:
payload["instrumental_id"] = args.instrumental_id
if args.n:
payload["n"] = args.n
print("Submitting instrumental generation task...", file=sys.stderr)
task = post_json("/v1/instrumental/generate", payload)
task_id = task["id"]
print(f"Task ID: {task_id}", file=sys.stderr)
print("Polling for completion...", file=sys.stderr)
result = poll_task("/v1/instrumental/query", task_id,
interval=args.poll_interval, timeout=args.poll_timeout)
download_choices(result, args)
def cmd_lyrics(args):
"""Generate or extend lyrics."""
if args.lyrics_command == "generate":
result = post_json("/v1/lyrics/generate", {"prompt": args.prompt})
title = result.get("title", "")
lyrics = result.get("lyrics", "")
if title:
print(f"Title: {title}\n")
print(lyrics)
elif args.lyrics_command == "extend":
result = post_json("/v1/lyrics/extend", {"lyrics": args.lyrics})
print(result.get("lyrics", ""))
def cmd_upload(args):
"""Upload a file to Mureka."""
print(f"Uploading {args.file} (purpose={args.purpose})...", file=sys.stderr)
result = upload_file_api(args.file, args.purpose)
file_id = result.get("id", "")
print(f"File ID: {file_id}")
print(json.dumps(result, indent=2))
# ---------------------------------------------------------------------------
# Shared argument helpers
# ---------------------------------------------------------------------------
def add_generation_args(parser, default_output="./output"):
"""Add common generation arguments."""
parser.add_argument("--model", default="mureka-8",
help="Model (default: mureka-8)")
parser.add_argument("-n", "--n", type=int, default=None, dest="n",
help="Number of results to generate (default 2, max 3)")
parser.add_argument("--output", default=default_output,
help=f"Output directory (default: {default_output})")
parser.add_argument("--format", choices=["mp3", "flac", "wav"], default="mp3",
help="Download format (default: mp3)")
parser.add_argument("--poll-interval", type=int, default=5,
help="Poll interval in seconds (default: 5)")
parser.add_argument("--poll-timeout", type=int, default=600,
help="Poll timeout in seconds (default: 600)")
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(
description="Mureka AI music generation CLI",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""examples:
%(prog)s song --lyrics "[Verse]\\nHello world" --prompt "pop, 120 BPM, female vocal"
%(prog)s instrumental --prompt "ambient, 80 BPM, soft pads"
%(prog)s lyrics generate "a summer love song"
%(prog)s lyrics extend "[Verse]\\nExisting lyrics..."
%(prog)s upload my_voice.mp3 --purpose vocal
""")
sub = parser.add_subparsers(dest="command", required=True)
# --- song ---
p_song = sub.add_parser("song", help="Generate a song with lyrics and vocals")
p_song.add_argument("--lyrics", required=True,
help="Song lyrics (max 3000 chars). Use structure tags: [Verse], [Chorus], etc.")
p_song.add_argument("--prompt", default=None,
help="Style/scene prompt (max 1024 chars)")
p_song.add_argument("--reference-id", default=None,
help="Reference audio file ID (purpose=reference)")
p_song.add_argument("--vocal-id", default=None,
help="Vocal file ID (purpose=vocal)")
p_song.add_argument("--melody-id", default=None,
help="Melody file ID (purpose=melody). Cannot combine with other control options")
add_generation_args(p_song, "./output")
# --- instrumental ---
p_inst = sub.add_parser("instrumental", help="Generate an instrumental track")
p_inst.add_argument("--prompt", default=None,
help="Style/scene prompt (max 1024 chars)")
p_inst.add_argument("--instrumental-id", default=None,
help="Reference instrumental file ID (purpose=instrumental)")
add_generation_args(p_inst, "./instrumental")
# --- lyrics ---
p_lyrics = sub.add_parser("lyrics", help="Generate or extend lyrics")
lyrics_sub = p_lyrics.add_subparsers(dest="lyrics_command", required=True)
p_lyrics_gen = lyrics_sub.add_parser("generate", help="Generate lyrics from a prompt")
p_lyrics_gen.add_argument("prompt", help="Theme or description for lyrics")
p_lyrics_ext = lyrics_sub.add_parser("extend", help="Extend existing lyrics")
p_lyrics_ext.add_argument("lyrics", help="Existing lyrics to continue writing")
# --- upload ---
p_upload = sub.add_parser("upload", help="Upload a file to Mureka")
p_upload.add_argument("file", help="Path to the audio file (mp3, m4a, or mid for melody)")
p_upload.add_argument("--purpose", required=True,
choices=["reference", "vocal", "melody", "instrumental", "voice", "audio"],
help="Purpose: reference (30s), vocal (15-30s), melody (5-60s), instrumental (30s), voice (5-15s), audio (general)")
args = parser.parse_args()
{"song": cmd_song,
"instrumental": cmd_instrumental,
"lyrics": cmd_lyrics,
"upload": cmd_upload}[args.command](args)
if __name__ == "__main__":
main()