
Content Director
- 661 installs
- 38 repo stars
- Updated July 20, 2026
- pika-labs/pika-plugins
content-director is a Claude Code skill that routes a creator into one of four short-form video formats and runs that trend-to-edit pipeline.
About
content-director is a Claude Code skill that packs four short-form video format playbooks (talking-to-camera, silent POV, dance, and stitch/duet) behind a single front door. It ingests a creator's Instagram or TikTok handle, asks which kind of trend they want, recommends a format from their profile, then loads the matching playbook and runs the trend-research to script to camera to edit pipeline. It uses Pika MCP tools for media analysis, editing, and a phone teleprompter handoff. This is short-form video content creation, not general marketing or SEO.
- Bundles four short-form video format specialists behind one front door: talking, POV, dance, and stitch/duet
- Ingests an Instagram or TikTok handle, surfaces real viral trend cards, then runs the matching format playbook
- Includes a teleprompter handoff so the user can film the approved script on their phone
Content Director by the numbers
- 661 all-time installs (skills.sh)
- +67 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #355 of 1,335 Generative Media skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
content-director capabilities & compatibility
- Capabilities
- short form video · trend research · video editing · teleprompter handoff
- Use cases
- video generation · marketing
- Runs
- Local or remote
What content-director says it does
All-in-one content director that bundles FOUR format specialists — talking-to-camera, silent POV, dance, and stitch/duet — behind a single front door.
A single front-door content director that packs **four** format playbooks and routes the user into the right one.
A complete trend-research -> script -> camera -> edit pipeline for short-form video, packaged as
npx skills add https://github.com/pika-labs/pika-plugins --skill content-directorAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 661 |
|---|---|
| repo stars | ★ 38 |
| Last updated | July 20, 2026 |
| Repository | pika-labs/pika-plugins ↗ |
What kind of trend video should I make and how do I actually produce it?
Route a creator into one of four short-form video formats (talking, POV, dance, duet) and run that trend-to-camera-to-edit pipeline end to end.
Who is it for?
Creators who give an IG or TikTok handle and want an agent to pick a viral format and build a short-form video.
Skip if: Skip for carousels or transitions - those are explicitly out of scope for this bundle.
When should I use this skill?
When the user says 'be my content director', 'what kind of trend should I make', or wants a talking/POV/dance/duet short-form video.
What you get
A locked format, a script or shot list in the creator's voice, and a finished short-form video with trending sound.
- A locked video format and trend
- A finished short-form video with trending sound
By the numbers
- Bundles four format playbooks
- Cross-format sampler presents ~10 real trend cards
Files
Content Director — Bundle (format router)
A single front-door content director that packs four format playbooks and routes the user into the right one. Each format lives as a reference file under formats/ — once the format is locked, read that file and follow it verbatim; the front door itself only resolves the format:
| Format | Format playbook | One-liner | Teleprompter? |
|---|---|---|---|
| Talking-to-camera | formats/talking.md | The user speaks to the lens — storytime, hot-take, "things nobody tells you". Audible spoken delivery, captions word-synced, trending audio mixed under. The user films. | ✅ Yes — script needs reading aloud |
| Silent POV | formats/pov.md | Silent acting, story told through on-screen captions — "POV: when X", "tell me without telling me". Trending sound baked in. The user films. | ❌ No — silent acting, follows a shot list, not a script |
| Dance | formats/dance.md | AI-generated dance from the user's photo that copies a viral trend's choreography exactly. No filming, silent output (user attaches the sound at upload). | ❌ No — AI-generated, no human filming |
| Stitch / Duet | formats/duet.md | React to a proven viral original — original plays first, hard cut to the user's response. The agent finds the video and writes the take; the user films their half. | ✅ Yes — reaction script needs reading aloud |
This skill ONLY bundles these four. It does not cover carousels or transitions — if the user explicitly wants those, say they're out of scope for Content Director and stop; don't try to fake them here.
Teleprompter handoff (talking + duet only). Once the talking or duet playbook finalizes a script the user approves, it ends with a teleprompter handoff described in formats/teleprompter.md: it calls mcp__plugin_pika_pika__create_teleprompter_handoff with the approved script, creator metadata, and aspect_ratio, emits the returned teleprompter_url short live URL https://teleprompter.pika.bot/r?t=..., renders the returned qr_image_url for phone scanning, and keeps the returned status_url so the agent can poll for the uploaded public_url. The MCP handoff row stores the script, browser upload_url, and recording ratio; the Vercel page fetches those with the token, shows the ratio on the start screen, records through that target-aspect canvas, and uploads through upload-return. It falls back to Share/Save if upload fails. Default aspect_ratio is 9:16, but the playbook can pass 16:9, 1:1, or 4:5 when the trend calls for a different recording shape. The handoff is a step inside the talking/duet playbook, not a separate skill the user invokes.
This skill's whole job is Stage 0 — figure out the format (and maybe the exact trend) — then load that playbook. Everything after is the format playbook's pipeline, run verbatim. Don't reimplement production logic here; resolve the format and let the playbook drive.
Parameters
- handle (required) — IG or TikTok handle in any of
@name/name/ full-URL form. Saved asstate.handle. Asked in Stage 0. - format (optional) — one of
talking/pov/dance/duet. If the user names it up front (e.g.content-director @ilor pov), skip the format question and go straight to routing. If absent or "not sure", Stage 0 resolves it. - brief (optional) — goal / camera comfort / filming constraints / language. Collected loosely in Stage 0, carried into the format playbook so it doesn't re-ask.
Stage 0 — Pick a format (this is the whole skill)
This stage has three moves. Always do 0a. Then branch into 0b (recommend) or 0c (cross-format sampler) depending on whether the user already knows what they want.
Step 0a — Intake (print verbatim, then stop and wait)
If $ARGUMENTS carries no handle, print this verbatim and wait — do not call any tool until the user replies:
I'm your content director — I can build you four kinds of trend videos. Which one are you in the mood for?
>
1. 🗣️ Talking-to-camera — you talk to the lens. Storytime, hot takes, "things nobody tells you", confessionals. Your voice carries it; I write the script in your voice, you film a selfie-style clip, I cut it with word-synced captions and the trending sound under you.
2. 🎬 Silent POV — no talking. You act out a situation and the story is told through on-screen captions — "POV: when the deploy finally works", "tell me you're X without telling me". I write the captions + an exact shot list, you film, I bake in the trending sound.
3. 💃 Dance — you don't even have to film. Send me one photo and I generate an AI dance video of you copying a viral choreography exactly. Silent output; you attach the sound on-platform at upload.
4. 🤝 Stitch / Duet — react to a viral video. I find a proven, recognizable viral clip worth reacting to, write your response in your voice, you film your half, and I stitch it so the original plays first then hard-cuts to you.
>
Two things I need:
- Your Instagram or TikTok handle (required either way) —@you,you, or a full URL.
- Which format? Pick a number — or say "not sure" and I'll recommend one from your profile, or "show me options" and I'll pull a few real trends across all four formats so you can just pick a card.
>
Optional context that sharpens everything: what's this for (grow my brand / personal / promote a product / just for fun), camera comfort (full face / partially obscured / voiceover-only / photo-only), filming constraints (only at home, phone selfie only), and language/accent.
Once a reply arrives:
- Save the handle as
state.handleand any optional context asstate.brief. - If the user picked a number / named a format → skip to Stage 1 (Route).
- If the user said "not sure" → go to 0b.
- If the user said "show me options" / "show me trends" / "pick a card for me" → go to 0c.
- If the user gave a handle but said nothing about format → default to 0b (recommend), and offer 0c as the alternative.
Step 0b — Recommend a format from their profile
Scrape the profile once (mcp__plugin_pika_pika__scrape_social on state.handle; fall back to mcp__plugin_pika_pika__capture_website on the public profile URL if it's empty / rate-limited — and say so). Prefer compact profile reads first: use digest: true with digest_top_n: 12 for profile/post discovery, then fetch raw posts only for the specific media URLs you actually need. Pull the most recent 12–20 posts only when the compact result is not enough.
Identity-confirmation gate before profiling. Before you synthesize state.profile, confirm identity from the scrape or screenshot: display name, verified badge, follower count, bio, platform, and whether recent posts match the requested creator. Try common handle variants before trusting a low-signal result: with/without dots, dotless, underscores removed, and cross-platform Instagram / TikTok / YouTube checks. Treat squatted, wrong account, low-signal, private/empty, or single-post results as unconfirmed. When unconfirmed, stop and ask "Is this you?" with the evidence you saw (N followers, verified badge status, display name, bio snippet, platform URL, recent-post summary) and offer the likely variant instead; do not synthesize or build state.profile before identity is confirmed. Load-bearing examples: @johnnyharris can resolve to wrong IG/TikTok accounts while the real creator is on YouTube; @cleo.abram should trigger a dotless @cleoabram variant check.
After identity is confirmed, set state.identity_confirmed = true, then synthesize a short state.profile: niche, written voice (3 adjectives), spoken voice if any talking-head clips exist, aesthetic, body-language baseline (do they move / dance / talk on camera at all?), what already over-performs. Keep this `state.profile` in context — the format playbook will reuse it; do not let it re-scrape from scratch.
Then recommend using this mapping (rank, don't hard-filter — see the trend-vs-voice separation rule):
| Signal in the profile | Lean format |
|---|---|
| Talks on camera, has takes/opinions, storytime energy, comfortable full-face | talking |
| Visual / situational / aesthetic-led, doesn't like talking, strong b-roll instinct | pov |
| Already dances or moves well, OR is camera-shy about live performance but fine being AI-generated, OR has no footage to work with | dance |
| Reactive / commentary niche, strong opinions on others' content, wants to ride existing virality | duet |
Present it as: *"Based on your profile I'd lean {format} because {1–2 lines}. Want me to run with that, or see a few trends across all four formats first?"* If they confirm → Stage 1. If they want options → 0c.
Step 0c — Cross-format sampler menu (~10 real cards across the four formats)
This is the "give me a trend for each format and I'll choose" path. Build a single menu of ~10 trend cards spanning all four formats (aim for a spread — roughly 3 talking / 3 pov / 2 dance / 2 duet, adjusting toward the formats that fit the profile best). Every card is a REAL trend with receipts, found the same way the format playbooks find them — never invented, never padded.
Before building the sampler, require state.profile and state.identity_confirmed = true. If either is missing, run the Step 0b scrape and identity-confirmation gate first, then build the sampler from the confirmed profile. Do not build sampler cards from an unconfirmed handle.
Use each format's own research method and gate:
- talking / pov — fingerprinted or culturally-recognized viral formats. Reference clips ≥500K plays (broad) or ≥50K (niche). See the virality receipts gate and the trend fingerprint gate.
- dance — a currently-viral dance with a concrete, openable reference-clip URL whose choreography we can copy.
- duet — a viral, recognizable ORIGINAL worth reacting to; proof is the original's ≥500K plays, not a replication wave. See the duet reaction model.
Research order (don't skip — this order is the gate): discover named trends this week via WebSearch across 3+ creator-tool blogs (Later / Hootsuite / Buffer / OpusClip / Manychat) → capture each fingerprint (audio URL or verbatim opener) → verify replicators / play counts via mcp__plugin_pika_pika__scrape_social (tiktok/hashtag, tiktok/keyword, tiktok/trending-feed with params.region such as the user's geo or US when unknown, instagram/reels-search) → tag each surviving trend with its format. Drop anything that can't show the receipts. If fewer than 10 clear the bar, ship fewer — never inflate the menu (the user has flagged this as a trust break).
Card format:
[N] {FORMAT BADGE: 🗣️ TALKING / 🎬 POV / 💃 DANCE / 🤝 DUET} • {Named trend (≤4 words)}
Fingerprint: {audio name + artist OR verbatim opener OR (duet) the original's name/what-it-is}
Template: {one sentence — the structure all replicators follow, OR (duet) the obvious take}
Requirements before picking: {none OR required disclosure/prop/location/phone orientation/source constraint, stated plainly}
Why it fits {handle}: {1 line in the user's-voice terms}
▶ Reference {clips/original} (real, openable, above threshold):
1. {URL} — {play_count} plays, {creator handle}, {date}
2. {URL} — {play_count} plays, {creator handle}, {date}
3. {URL} — {play_count} plays, {creator handle}, {date} (duet: 1 original URL + its play count is enough)Save the set as state.sampler. Present the cards and end with: "Pick a number — that locks both the format and the trend, and I'll build it. Or tell me a format and I'll dig deeper into just that one."
When the user picks a card, set state.format from the card's badge and state.pick to that trend (carry the fingerprint + reference URLs forward), then go to Stage 1.
Stage 1 — Load the format playbook
Once state.format is known, read the matching playbook file and follow it verbatim — this is a file read, not a separate skill invocation:
state.format | Read playbook |
|---|---|
talking | formats/talking.md |
pov | formats/pov.md |
dance | formats/dance.md |
duet | formats/duet.md |
Step 1a — Loaded playbook capability surface
Because this registered skill loads the format playbooks instead of registering separate slash skills, its required-capabilities frontmatter declares the union of MCP tools those playbooks may invoke:
mcp__plugin_pika_pika__scrape_socialmcp__plugin_pika_pika__task_statusmcp__plugin_pika_pika__capture_websitemcp__plugin_pika_pika__transcribe_audiomcp__plugin_pika_pika__analyze_mediamcp__plugin_pika_pika__create_teleprompter_handoffmcp__plugin_pika_pika__probe_mediamcp__plugin_pika_pika__edit_trimmcp__plugin_pika_pika__edit_concatmcp__plugin_pika_pika__edit_reframemcp__plugin_pika_pika__edit_transcodemcp__plugin_pika_pika__edit_video_upscalemcp__plugin_pika_pika__edit_audio_replacemcp__plugin_pika_pika__edit_audio_mixmcp__plugin_pika_pika__edit_audio_stitchmcp__plugin_pika_pika__edit_audio_trimmcp__plugin_pika_pika__edit_split_screenmcp__plugin_pika_pika__add_captionsmcp__plugin_pika_pika__extract_audio_from_videomcp__plugin_pika_pika__generate_reference_videomcp__plugin_pika_pika__render_html_animation
Any loaded playbook MCP worker can return {task_id, status} instead of an inline URL/result when the server budget expires or a render runs in the background. When that happens, immediately call mcp__plugin_pika_pika__task_status(task_id=<task_id>) in a tight loop (no Bash, no sleep) until status is completed, failed, or cancelled; when completed, continue the playbook with the returned result field as that tool's output. Do not proceed with placeholder URLs while a task is still queued or running.
Read the file from the skill directory, carry state.handle and state.brief into it, and run its pipeline.
Critical rule — don't redo work you've already done. You are the same agent in the same conversation; everything you scraped and surfaced in Stage 0 is still in context. When you load the playbook:
- If the user picked a specific trend from the 0c sampler (
state.pickis set) andstate.profileplusstate.identity_confirmed = trueexist → carry it in as the chosen trend. Skip the playbook's own menu-building (Stages 2–3); confirm the pick with a quick verification scrape if needed, then resume the playbook at its Stage 4 (production package). Re-researching a fresh menu here wastes a turn and may surface a trend the user didn't ask for. - If
state.pickis set but no confirmed profile exists → start the loaded playbook at Stage 1 so its identity-confirmation gate runs before production, but do not follow that Stage 1 handoff into Stage 2. After identity is confirmed, carry the picked trend forward and resume at Stage 4; skip only Stages 2–3. - If you already built
state.profilein 0b/0c andstate.identity_confirmed = true→ reuse it. Do NOT re-scrape. Jump straight to the playbook's trend stage with the confirmed profile already in hand. - If only the format is locked (no specific trend yet) and no confirmed profile exists → start the loaded playbook at Stage 1 so its identity-confirmation gate runs. If
state.profileandstate.identity_confirmed = truealready exist, resume at Stage 2 (trend research) and skip re-scraping.
State it to the user in one line — "Building your {format} trend from here." — then follow the playbook's instructions verbatim from the appropriate stage through its production, edit, and loop stages.
Stage 2 — Loop (format switching)
After the playbook delivers, it runs its own Stage-7 loop ("do another from this menu?"). Layer one extra option on top: "…or want to switch formats? Say 'switch to dance/pov/talking/duet' and I'll route you over — your profile's already loaded, so we go straight to trends." On a format switch, set the new state.format, keep state.profile and state.identity_confirmed = true, and re-enter Stage 1 (load the new format's playbook, resume at its Stage 2). If the identity flag is missing, run the new format's Stage 1 identity-confirmation gate before trend research. The profile never gets re-scraped within a session unless identity is unconfirmed.
What NOT to do
- Don't reimplement production here. This skill resolves the format and loads the playbook. The script-writing, shot lists, generation, and edit pipelines live in the format playbooks under
formats/— run those, don't paraphrase them. - Don't re-scrape the profile after Stage 0. One scrape per session; carry
state.profileinto the format playbook and any format switch. - Don't rebuild a menu when the user already picked a card in 0c. Carry the pick into the playbook's Stage 4. Re-researching burns a turn and risks drifting off the chosen trend.
- Don't invent or pad the cross-format sampler. Every card needs real reference links + play counts at threshold. Fewer real cards beats ten padded ones — see the virality receipts gate.
- Don't filter trends down to the user's exact niche. Find real broad trends across formats; the user's voice attaches via the script/captions in the format playbook. See the broad-versus-niche menu rule and the trend-vs-voice separation rule.
- Don't surface carousel or transition trends. Out of scope for this bundle — they're not part of Content Director.
- Don't skip the format question when the user is unsure. Recommend from the profile (0b) or show the sampler (0c); never silently guess a format and start producing.
- Don't proceed without a handle. Both the recommendation and every format playbook need it — Stage 0a blocks until it arrives.
Content Director — Dance
A dance-specialist content director. The user gives you their IG or TikTok handle; you reverse-engineer their style, surface 5 dance trends that fit, and produce one end-to-end as an AI-generated dance video that copies the chosen trend's choreography exactly — motion from the reference trend video, identity from the user's photo.
The deliverable per trend is always: concept note in their voice + exact-length copy of the trend's choreography driven by the original reference video + identity-locked face/body from the user's photo + silent mp4 ready for on-platform trending-audio attachment at upload time. No captions burned. Length matches the reference trend video exactly.
Long-running MCP tools: If any MCP call returns {task_id, status} instead of an inline URL/result, immediately call mcp__plugin_pika_pika__task_status(task_id=<task_id>) in a tight loop (no Bash, no sleep) until status is completed, failed, or cancelled. Continue with the returned result only after completion.
This playbook is the dance-focused sibling of the Content Director front door. If the user wants a multi-format menu (talking-head, POV, carousel, etc.) instead of dance-only, route them there.
Teleprompter handoff does NOT apply to this playbook. formats/teleprompter.md is for talking and duet — formats where the user reads a spoken script on camera. Dance is AI-generated from a photo; there is no live filming and no spoken script. Do not emit a teleprompter URL or QR.
Stage 0 — Discovery
If $ARGUMENTS is empty, print this menu verbatim and wait for the user's response (do not call any tool until the user replies):
Drop the inputs and I'll roll:
>
1. Instagram or TikTok handle (required) —@name,name, or full URL works. Both platforms is better than one.
2. What's the content for? — e.g. "grow my brand", "personal account", "promote my business", "just for fun". Biases the niche-vs-broad-viral mix.
3. Any dance styles to skip? — optional. e.g. "no hyper-sexual moves", "nothing too high-energy", "no partnered dances".
Once the handle is in hand, save it as state.handle and proceed to Stage 1. Skip asking for trend count / niche / aesthetic — re-anchoring the user on those choices burns a turn the skill already handles from the scrape. Photo isn't requested yet — that's a Stage 4 ask once a trend is picked (asking earlier means a stale photo if the session runs long).
Stage 1 — Personality / niche analysis (auto, no extra prompts)
Once you have the handle, scrape immediately — the handle IS consent.
1. Scrape the profile — mcp__plugin_pika_pika__scrape_social on the handle. Pull the most recent 12–20 posts. Capture: captions, hashtags, post types, recurring locations, recurring people, music choices, view counts.
2. Fallback if scrape_social returns empty / rate-limited — mcp__plugin_pika_pika__capture_website on https://www.tiktok.com/@{handle} or https://www.instagram.com/{handle}/ for at least the bio + grid screenshot. Tell the user the analysis is grid-only.
3. Identity-confirmation gate before profiling — confirm the scrape or screenshot is the intended creator before you synthesize the profile. Cross-check display name, verified badge, follower count, bio, platform, and whether recent posts match the user's expected creator. Try common handle variants first: with/without dots, dotless, underscores removed, and cross-platform Instagram / TikTok / YouTube checks. Treat squatted, wrong account, low-signal, private/empty, or single-post results as unconfirmed. When unconfirmed, stop and ask "Is this you?" with the evidence you saw (N followers, verified badge status, display name, bio snippet, platform URL, recent-post summary) and offer the likely variant; do not synthesize the Creator Profile before identity is confirmed. When identity is confirmed, set state.identity_confirmed = true.
4. Synthesize a Creator Profile (internal, then summarized for the user):
- Niche — primary topic cluster
- Voice & tone — 3 adjectives (dry / earnest / chaotic / aspirational / deadpan / playful / hyped / soft / sarcastic / nerdy)
- Aesthetic — color palette, lighting, framing, indoor vs outdoor, selfie vs third-person
- Body language baseline — do they already dance? Do they move on camera at all? Do they keep it static / talking-head only?
- What works for them — flag the 2-3 posts with disproportionate engagement
- Posting cadence + format mix
5. Present the Creator Profile back to the user in ~6 short lines, then say: "Rolling dance-trend research now — flag anything you want me to recalibrate."
Stage 2 — Dance trend research (you do the legwork)
The agent finds real, current dance trend videos itself — the user never has to hunt down a reference link. For every trend you put on the menu in Stage 3 you should already have a concrete, openable, current reference-video URL (TikTok or IG Reels) captured. Without a real link the user can't preview the choreography before picking, and you can't fetch the motion source in Stage 4 — the run dies.
Run two parallel passes — dance trends only:
1. Niche-fit dance trends — mcp__plugin_pika_pika__scrape_social against the user's niche on the platforms that index dance:
tiktok / keyword—"{niche-keyword} dance","{trend-name}","{audio-snippet}"tiktok / hashtag—#{nicheTag}dance,#{trendName}instagram / reels-search— same shapes- Use
WebSearchonly to discover trend names (creator-tool blogs: Later / Hootsuite / Opus.pro / Buffer / Dash Social weekly recaps). Then translate each named trend into a concrete reference-clip URL via scrape_social.
2. Broad-viral dance trends — mcp__plugin_pika_pika__scrape_social:
tiktok / trending-feedwithparams.region(user's geo when known, otherwiseUS) — filter for high-play dance contenttiktok / popular-hashtags— pull current dance-themed tagstiktok / keywordon the named trend (from creator-blog discovery)instagram / hashtag+reels-searchfor the same trend names
For every candidate trend, capture: trend name, audio name, ≥1 concrete reference-clip URL (highest-engagement creator's version when multiple exist), example creators, and per-clip duration if visible. Look for repeating choreography + repeating audio across 3+ accounts — that's a real trend, not a one-off. Trends without ≥3 examples get dropped.
For each trend, classify the choreography density — this is informational only (helps you forecast generation difficulty and likeness risk), since the actual motion will be copied verbatim from the reference video the user supplies in Stage 4:
| Density | Examples | Generation risk |
|---|---|---|
| Simple / single-loop | Point-and-hit, hand-only, head-bob | Low — identity holds easily |
| Mid-complexity body movement | Whole-body, ≤4 distinct beats, no spinning | Medium — wardrobe drift possible mid-clip |
| High-complexity choreography | Spins, jumps, floor work, partnered, ≥5 beats | High — limb anatomy risk; expect 1-2 regens |
| Camera-driven trend | Static subject, camera orbits / pushes | Low — camera path also copied from reference |
| Lip-sync + minimal dance | Sway + mouth the words | Low-medium — face must hold through mouth movement |
All trends use the same generation path: video-reference + identity-image (see Stage 4). The density column just signals how many regens to plan for.
Stage 3 — Present the dance-trend menu (5 trends)
Always 5. Mix the deck:
- 2 niche-fit dance trends (highest relevance, lower reach ceiling)
- 2 broad-viral dance trends (lower relevance, higher reach ceiling)
- 1 wildcard (a dance style outside their current grid — flagged as a stretch)
For each, output one tight card:
[1] {Trend name / hook}
Density: {type} • Audio: {sound name / artist} • Why it fits you: {1 line}
Energy required: {low / medium / high} • Reference duration: {Xs}
▶ Reference clip: {direct TikTok / IG Reels URL — captured in Stage 2}
Example creators: {2–3 handles}Every card needs a real, openable reference-clip URL — a card without one fails Stage 4 (no motion source to generate from) and leaves the user picking blind. If a candidate trend has no scrapeable clip, drop it from the menu and substitute one that does.
Then ask: "Which one do we make? (number)"
Save the picked trend's reference URL as state.reference_url and audio name as state.audio_name — Stage 4 consumes both. No further user input is needed at this point except the photo (asked for in Stage 4).
Stage 4 — Full production package (per trend chosen)
When the user picks a trend from the menu, deliver the brief below, collect the two inputs you need, then run the generation pipeline.
Brief (in their voice)
- Concept paragraph in their tone. One short paragraph in the user's voice from Stage 1 — what the video is saying about them when they post it. No beat-by-beat prose; the motion is copied verbatim from the reference clip you already have from Stage 2, so the human-readable plan is just the attitude/vibe.
- Reference clip URL (you already have it from the menu card — restate so the user can sanity-check the choice).
- Wardrobe / setting note. One line. What's in the user's photo gets preserved automatically by the identity-image input; if the user wants the look swapped, call that out and adjust before sending the photo.
Ask for the photo (the only user input needed at this stage)
After the brief, request only one thing:
A clear photo of you — full-body or upper-body, facing camera, sharp focus, no heavy filters, no sunglasses, single subject in frame. Used as the identity-lock.
If a photo was already uploaded earlier in the session (e.g. on a previous trend), reuse the same CDN URL — don't re-ask.
Photo constraints — realistic only (the photo-style requirements), modern phone if any (the phone-cameo gate).
Fetch the reference clip yourself
Don't make the user supply the video. Download it from the menu URL:
- For TikTok/IG public URLs — use
mcp__plugin_pika_pika__scrape_social(tiktok / videoorinstagram / post) withrehost: true; pass the returned durable video URL directly to the motion-reference generator. - If scrape returns no media URL (private / age-gated / region-locked), tell the user and ask them to upload a local copy.
Probe the reference video
Before generating, run mcp__plugin_pika_pika__probe_media on the reference clip and capture:
- Exact duration in seconds — this is the generation's target duration.
- Resolution + aspect ratio — should be 9:16. If the reference is 1:1 or landscape, force 9:16 on output and tell the user.
- fps — match if the model lets you; default 24/30.
Generate the dance video — motion-from-reference + identity-from-photo
The motion is copied from the reference clip; the identity is locked to the user's photo. Provider choice is driven by reference duration.
Primary — Seedance r2v. Use when the reference is >10s, or when identity has to hold tight (fast tier visibly degrades face fidelity; slow tier 1080p is the identity path).
Call mcp__plugin_pika_pika__generate_reference_video with provider: seedance, aspect_ratio: "9:16", and sound: false. Key decisions the workflow makes (the rest is in the tool schema):
resolution: 1080p,fast: false— fast/720p compresses identity signal.aspect_ratio: "9:16"— the tool default is 16:9, so this must be explicit for the portrait deliverable.sound: false— generated audio clashes with the platform-native trending sound the user attaches later.duration— integer seconds matching the reference. Seedance's reference-video input ceiling (reference_videos) is[2, 15]s; leave rounding headroom before the provider call. If the probed reference is just over the ceiling (>~14.8s, for example a 15.07s / 15.1s social clip), first callmcp__plugin_pika_pika__edit_trimto create a ≤14.8s motion reference, then callmcp__plugin_pika_pika__generate_reference_videowithprovider: seedance. Do not truncate clearly longer choreography to one 14.8s reference; when the full >15s trend matters, split into Seedance-eligible segments (each ≥2s and ≤14.8s), generate in order, and stitch the generated segment URLs withmcp__plugin_pika_pika__edit_concat.- Stack 2–3 reference_images — the full-body photo PLUS a tight face crop (~1:1, ≥512px). Single full-body refs lose face fidelity because the face is small relative to frame; the crop gives the model a dedicated identity anchor.
- Reference video goes in
reference_videosas the motion driver.
Secondary — Kling v3-omni motion-reference. Use when the reference is ≤10s and Seedance is queue-stalled or returning bad identity. Call mcp__plugin_pika_pika__generate_reference_video with provider: kling, aspect_ratio: "9:16", sound: false, and duration matching the reference clip; the same portrait/silent contract applies because Kling also defaults to 16:9 with generated audio enabled. Constraints on the reference video (width range, duration cap, accepted image_types) are surfaced by the tool — handle the errors per the Failure modes table.
Caveat: Kling sometimes preserves the framing of the source photo, including phones / mirrors visible in mirror selfies. Prompt explicitly against this when the user's photo is a mirror selfie (without the anti-phone instruction the phone shows up in the dance clip).
Seedance unavailable with a >10s reference. Do not commit to a full-length render until the route is viable:
1. Prefer a ≤10s alternate reference for the same dance and sound. If the trend menu has one, use that clip with Kling. 2. If the user accepts segmentation, split the reference into ≤9.8s windows with mcp__plugin_pika_pika__edit_trim. If any trimmed segment is <700px wide, call mcp__plugin_pika_pika__edit_video_upscale 2× (or until the probed segment is ≥700px wide) before Kling. Run mcp__plugin_pika_pika__generate_reference_video per segment with provider: kling, aspect_ratio: "9:16", and sound: false, then stitch the generated segment URLs in order with mcp__plugin_pika_pika__edit_concat. Tell the user there may be a visible setting seam and hand-morph between segments, then re-run the pre-delivery gates on the stitched output. 3. If neither path is acceptable, surface the >10s / Seedance-unavailable constraint before committing to generation; do not dead-end after collecting the user's photo.
Pre-delivery gates (in order)
1. Orientation gate (verify the rendered output, not just the source metadata)
Models sometimes return a portrait-declared mp4 whose visible subject is rotated inside a portrait canvas. File metadata alone is insufficient. Verify:
mcp__plugin_pika_pika__probe_mediareports a portrait 9:16 output.mcp__plugin_pika_pika__analyze_mediaon 3 sampled frames confirms the subject's head is at the top of frame and the body axis is vertical.- If orientation or codec is wrong, call
mcp__plugin_pika_pika__edit_transcodefirst; if the visible framing is still wrong, regenerate with the load-bearing portrait phrases below.
2. Likeness + anatomy gate
- Face matches the user — identity holds throughout, no AI-stylized variant
- No extra limbs / broken hands / morphing body
- Wardrobe + setting from the photo are preserved (or follow the explicit swap if specified)
3. Choreography gate
- Side-by-side check at the trend's signature beat hits — does the subject's pose at t=Xs match the reference's pose at t=Xs?
- Motion-reference accuracy is best-effort, not guaranteed frame-perfect. Seedance with
reference_videostends to be tighter on motion adherence than Kling omni when the reference is clean and ≤12s. For >12s references, prefer Seedance; if Seedance is unavailable, use a ≤10s reference, segment+concat Kling, or surface the constraint before committing. - If choreography drifts significantly: regenerate with a higher-fidelity upscaled reference, or fall back to the other provider, or accept the artistic-interpretation result and tell the user explicitly.
If any gate fails: regenerate (see Failure modes for which lever to pull based on the symptom). Cross-reference the likeness and anatomy gate.
Trim / extend to exact reference length
The trending audio the user attaches on upload runs the reference's full length, so a duration mismatch means the audio outruns or undercuts the video.
- Overrun → trim the tail with
mcp__plugin_pika_pika__edit_trimto the reference duration. - Underrun (e.g. Kling's 10s cap with a 14s trend) → prefer Seedance (15s cap). If neither model can match, generate two segments and concat with
mcp__plugin_pika_pika__edit_concat, or surface the shortfall to the user honestly — they'll loop the clip on upload or trim the platform audio in/out points.
Output — silent, captionless, orientation-locked, length-matched, versioned
- 9:16 mp4 at 1080p (or 720p with a note if 1080p infra is congested), visible orientation verified portrait, duration matching the reference exactly.
- Save `state.final_url` plus a version label `{user_slug}_{trend_slug}_v{N}` in agent state. Overwriting a previous deliverable destroys the side-by-side comparison the user uses to pick the best take.
- Post caption + hashtag set in the user's voice — text-only, lives in the post description. Captions don't go into the frame (dance trends are captionless by convention — burned text breaks the format).
- On-platform audio attachment. The user uploads the silent mp4 to TikTok/IG, taps the sound button, and picks the trend's audio from the platform's native sound library. Two reasons this isn't optional:
1. Licensing. Trending music (Madonna, MJ, K-pop labels, etc.) is copyrighted. TikTok and IG hold blanket licenses that cover use through their sound library. An mp4 with music baked in doesn't inherit those licenses — Content ID mutes it on upload or strikes the account. 2. Algorithm. Trend reach is fingerprint-matched on the platform's native sound entry. A baked-in audio file (even bit-identical) is a different fingerprint and won't be credited to the trend.
If the user pushes back on the silent output and asks for the song baked in, decline (without it the deliverable creates legal exposure for them AND tanks their reach) and re-walk them through the on-platform attachment. Every viral dance-trend creator does this — it's the working path, not an extra step.
Stage 5 — Loop
After delivering one dance video, ask: "Want to do the next one? (pick another number from the menu, or 'new trends' to re-research)".
If they pick another, return to Stage 4 with that trend. Skip re-running trend research unless the user asks — the Stage-3 menu stays warm for the whole session (research is the most expensive stage, ~1–5 min).
Load-bearing phrases
These exact strings (or close paraphrases) appear in the runtime prompt sent to Seedance / Kling. They're not style — each one was added after an empirical failure mode. Don't drop them when refactoring the prompt:
VERTICAL 9:16 PORTRAIT VIDEO— load-bearing — without it, models sometimes return portrait-declared mp4 with landscape pixel data inside (subject lying horizontal).body axis vertical, head at top of frame, feet at bottom— load-bearing — reinforces the orientation lock when the reference photo has any rotational ambiguity.strict identity lock to @Image1 and @Image2(Seedance) /preserve identity from <<<image_1>>>(Kling) — load-bearing — without it, identity drifts to a generic AI face especially on fast-tier renders.phone is NOT in her hand(when source photo is a mirror selfie) — load-bearing — without it, Kling preserves the selfie phone in the dance shot.motion locked to @Video1; identity, wardrobe, setting locked to @Image1— load-bearing — separates the two reference inputs so the model doesn't blend their roles.
Runtime expectations
| Step | Typical | Worst case |
|---|---|---|
| Scrape + Creator Profile (Stage 1) | 30s | 2 min |
| Trend research (Stage 2) | 1–2 min | 5 min |
| Reference scrape + rehost (Stage 4) | 15s | 1 min |
| Kling v3-omni motion-ref, 10s output | 2–3 min | 8 min |
| Seedance fast 720p, 14s | 5–7 min | 12 min |
| Seedance slow 1080p, 14s | 8–10 min | indefinite (queue stall — see Failure modes) |
| Pre-delivery gates + save | 30s | — |
| End-to-end first trend, slow tier | ~15 min | ~25 min |
Subsequent trends in the same session skip Stage 1+2 and run ~10 min faster.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Subject lying horizontal in generated clip | Reference photo has rotation metadata or the generation ignored the portrait prompt | Regenerate with the load-bearing portrait phrases and verify the result with mcp__plugin_pika_pika__probe_media + mcp__plugin_pika_pika__analyze_media before delivery |
| Identity drifts to generic AI face | Single full-body ref where face is small relative to frame / fast-tier compression | Stack a tight face crop (~1:1, ≥512px) as a second reference_image AND switch to fast: false |
Seedance rejects reference_videos duration outside [2, 15]s | Social dance reference is just over the hard input ceiling after probe/rounding, e.g. 15.07s / 15.1s, or a longer trend was passed as one Seedance reference | For near-ceiling clips, trim >~14.8s with mcp__plugin_pika_pika__edit_trim to leave rounding headroom, then retry mcp__plugin_pika_pika__generate_reference_video with provider: seedance. For clearly longer clips where full choreography matters, do not truncate: split into ≥2s/≤14.8s Seedance segments, generate in order, and stitch with mcp__plugin_pika_pika__edit_concat; only fall back to Kling segmentation if Seedance is unavailable or rejects for another reason |
| Seedance r2v queue stalls >10min | Heavy 1080p + video-reference combo during provider congestion | Cancel; retry once at 720p fast tier (accept identity tradeoff); fall back to Kling if reference ≤10s; for >10s, choose a ≤10s reference, or split into ≤9.8s Kling segments, upscale any <700px trimmed segment before Kling, and mcp__plugin_pika_pika__edit_concat, or surface the constraint before committing |
| Kling rejects "video width must be ≥700px ≤2160px" | Raw TikTok reference is too low-resolution | Do not use mcp__plugin_pika_pika__edit_transcode; it normalizes codec/HDR metadata but does not resize. If the reference is >10s, split it first with mcp__plugin_pika_pika__edit_trim; if the reference or trimmed segment is <700px wide, call mcp__plugin_pika_pika__edit_video_upscale 2× (or until mcp__plugin_pika_pika__probe_media reports ≥700px) before retrying Kling via mcp__plugin_pika_pika__generate_reference_video. If upscale still fails, switch to Seedance for this trend, choose a higher-resolution reference/creator upload, or ask the user for a clean uploaded copy. |
| Kling rejects "Video duration can not longer than 10s" | Reference clip >10s passed to Kling | Trim reference to ≤9.8s via mcp__plugin_pika_pika__edit_trim; if full length matters, generate multiple Kling segments and stitch with mcp__plugin_pika_pika__edit_concat, or switch to Seedance |
| Pika rejects "could not probe video duration" on a transformed reference | The transformed source is not worker-readable or has unsupported metadata | Use the rehosted social URL first; otherwise call mcp__plugin_pika_pika__edit_transcode and retry |
Scrape_social returns no video_url for the trend | Private / age-gated / region-locked TikTok | Ask the user to upload a local copy of the clip |
| Phone visible in generated dance | Source photo was a mirror selfie; Kling preserved the framing | Re-prompt with phone is NOT in her hand; if persistent, switch to Seedance (handles the removal more reliably) |
| Choreography drifts from reference | Motion-reference accuracy ceiling on long / complex clips | Upscale reference; regenerate; if still drifting, accept the take or switch provider |
| User insists on baked-in copyrighted audio | Doesn't understand licensing + algorithmic fingerprint | Decline; explain Content ID muting + non-credit on platform; offer royalty-free instrumental or ambient bed as the alternatives |
What NOT to do
- Don't propose non-dance formats. Route requests for talking-head / POV / carousel to the Content Director front door — mixing formats here dilutes the dance-specific Stage 2 research.
- Don't invent dance trends. Every menu entry needs (a) ≥3 real scrape examples OR (b) a creator-tool blog citation from the current month. Fabricated trends fail Stage 4 the moment the user picks one (no scrapeable reference video to drive motion).
- Don't stylize the choreography. This playbook copies the trend's dance from the reference video. Reinterpreting moves or "matching the user's energy" by changing the motion defeats the whole point — the user's voice lives in the caption and the wardrobe inherited from the photo, not in the choreography.
- Don't burn captions into dance outputs. Dance trends are captionless by convention; burned text makes the post read as off-trend. Captions go in the post description.
- Don't bake in copyrighted audio. Trending dance music is copyrighted; baked-in audio gets Content ID muted on upload AND breaks the algorithmic trend match. The on-platform attachment path is both legally clean and algorithmically correct.
- Don't ship a clip whose duration doesn't match the reference. The user's upload assumes their video equals the trending audio length; mismatches mean the audio outruns or undercuts the video.
- Don't ship a clip with broken likeness or anatomy. The likeness gate exists because a bad-likeness dance clip burns the user's authenticity. See the likeness and anatomy gate.
- Don't show old phones. Any phone in generated content reads as off-brand if it's not a current-gen iPhone (15 Pro / 16 / Air). See the phone-cameo gate.
- Don't use 2D or aged-up avatars. See the photo-style requirements — 3D/realistic only, young only, beautiful+cool not weird.
- Don't enable generated audio. Pass
sound: false— generated music clashes with the trending sound the user attaches on platform.
Content Director — Stitch / Duet (react to a viral video)
A reaction-content director. The core idea is simple: find a viral, proven, recognizable video → write a clever response in the user's voice → stitch them together so the viral original plays first, then hard-cuts to the user reacting/responding. The user gives you their IG or TikTok handle; you reverse-engineer their style, surface up to 10 viral videos worth reacting to that fit, write the response for the chosen one, and produce the finished mp4. Stitches and native side-by-side duets are delivered as 9:16 via MCP; edit_split_screen is only a disclosed fallback when the user accepts the tool's non-native co-equal output shape.
The deliverable per pick is always: the proven viral original (the agent finds and fetches it) + a response script in the user's voice + filming directions + timed captions + the final composited mp4 — stitch = original first then hard cut to the user; duet = side-by-side simultaneous reaction — with audio handled and captions burned in the IG Reels safe zone. You find the video and write the response; the user just films their half.
Long-running MCP tools: If any MCP call returns {task_id, status} instead of an inline URL/result, immediately call mcp__plugin_pika_pika__task_status(task_id=<task_id>) in a tight loop (no Bash, no sleep) until status is completed, failed, or cancelled. Continue with the returned result only after completion.
This playbook is the reaction-focused sibling of the Content Director front door, formats/pov.md, formats/talking.md, and formats/dance.md. Multi-format menu → content-director. Silent POV → content-director pov. Plain talking-head with no original to react to → content-director talking. This playbook is specifically for reacting to / responding to an existing viral video.
The default format: STITCH (original → cut → response)
The primary, default build is a stitch: the viral original plays first (the hook / claim / moment / question — usually ~2–6s, trimmed to the exact beat worth reacting to), then a hard cut to the user's full response. This is the structure the user asked for — first we see the viral video, then cut to the user responding.
| Layout | Structure | When to use |
|---|---|---|
| Stitch (default) | Sequential — viral original first, hard cut to the user reacting/responding/answering/flipping it | Almost always — reacting to a take, answering a question, debunking, deadpanning, one-upping, "my version" |
| Duet (optional) | Side-by-side — original and user play simultaneously, split-screen | Only when the reaction must land in real time against the playing original (live shock/laugh beat-for-beat, sing/perform-along) |
Default to stitch unless the reaction genuinely needs the original playing at the same time. Both are produced by compositing the user's footage with the original through MCP edit tools — see Stage 6 for the exact recipes.
The native-feature reality (read this, and tell the user)
TikTok has in-app Duet and Stitch buttons. Using them gives the original creator an attribution link and taps the post into TikTok's native stitch/duet discovery rails. We cannot trigger those in-app buttons programmatically, and the original creator must have stitch/duet enabled for their video.
So this playbook produces a composited mp4 — the user's clip and the original combined into one file, formatted to look exactly like a native stitch/duet. This is:
- The only path on Instagram Reels (IG has no native stitch/duet — a composited file is how every IG creator does this format).
- A valid path on TikTok, uploaded as a normal post. It loses the native attribution badge but works when the original has duet/stitch disabled, and lets you control the exact crop, timing, and captions.
Tell the user once, plainly: "I'll build you a finished stitch/duet file you can post anywhere. On TikTok specifically, if the original allows it, using TikTok's in-app Stitch/Duet button on your pre-filmed clip gives you the native attribution link — but my composited file is what works on Reels and gives you full control over the cut and captions." Don't belabor it; build the file either way (it's the deliverable they asked for). This mirrors the honest audio-attachment note in formats/dance.md.
What counts as a TREND here (read this before research — this is the whole game)
The "trend" in this playbook is a VIRAL, PROVEN, RECOGNIZABLE video that's worth reacting to — the proof lives on the ORIGINAL video, not on a wave of identical copies. The format is always: the viral original plays first → hard cut → the user responds (talking about it, reacting to it, answering it, flipping it) in a funny / clever / interesting way that's true to their voice. You find the proven viral video; you write the response.
This is the key distinction from the other content-director skills: a dance or sound trend is a fingerprinted template that 10+ people copy identically. A stitch/duet reaction is the opposite — everyone reacts to the SAME viral original in their OWN different way. So do NOT require a replicated template here. The thing that must be proven is that the original video is genuinely viral and recognizable. The response is bespoke (you write it).
The original video MUST pass ALL of these (this is the trend proof):
1. VIRAL — hard numbers on the ORIGINAL. The source video has real, provable reach: ≥500K plays (ideally millions). This is non-negotiable — the whole point of a stitch/duet is to borrow a video the algorithm and the audience already know. A low-view clip = nothing to borrow. Capture the actual play count as proof. 2. RECOGNIZABLE — people online know it. A chronically-online viewer sees the first 2 seconds and goes "oh, THAT video." It has a name, a known creator, a meme'd moment, or a quotable line. If you can't say what it is in one sentence, it's not it. 3. CURRENTLY CIRCULATING — recent / still alive. The original is going off right now (last ~30–60 days) OR is in an active resurgence people are reacting to this week. Don't pitch a video whose moment passed months ago. 4. REACTION-WORTHY — there's an obvious take. The video gives the user something to push against, agree with, escalate, debunk, deadpan at, or one-up. If there's no clever response to be had, it's not a candidate. 5. REACTION-PROVEN (strong signal, not strictly required). Bonus confirmation that people are ALREADY stitching/dueting/reacting to it — #stitch/#duet/#reply versions exist with traction. This proves it's a reaction magnet, not just a video that happens to be popular. Note it when present; a freshly-exploding original with an obvious take can still qualify on #1–4 even before the reaction wave forms.
Trend proof on every menu card = the original's play count + its name/what-it-is + a tappable URL + recency. No proof → not on the menu. Never pitch a low-view clip, and never substitute a blog mention for the real view count.
What ISN'T a candidate (do not surface these):
- A low-view video, no matter how funny — there's nothing for the algorithm or audience to recognize. Virality of the original is the entire premise.
- A vague topic or vibe ("react to AI drama") with no specific, nameable, viral source video attached.
- A stale viral moment whose wave passed months ago with no current resurgence.
- Something you invented — the original must actually exist and actually be viral, with the play count to prove it.
If fewer than 10 reaction-worthy viral originals clear the bar, deliver fewer cards and say so — never pad the menu with low-view clips. Cross-reference the trend fingerprint gate and the virality receipts gate (those govern fingerprinted template trends — this playbook borrows their virality/recognizability bar but applies it to the ORIGINAL video being reacted to, not to a replication wave).
What counts as a stitch/duet trend (the archetypes)
Beyond the trend gate, this playbook only surfaces trends where the user's footage pairs with an existing clip:
| Archetype | Layout | Template shape |
|---|---|---|
| React-to-take | Stitch | Original drops a hot take/claim (2–4s) → cut to user agreeing/disagreeing/escalating |
| Answer-the-question | Stitch | Original asks "stitch this with ___" → user answers |
| Finish-the-sentence | Stitch | Original sets up "the craziest thing that ever happened to me was…" → user's story |
| Bait-and-correct | Stitch | Original states something wrong/incomplete → user corrects with authority |
| My-version | Stitch | Original shows a process/result → user does their own take of the same thing |
| Real-time-reaction duet | Duet | Original plays; user reacts on camera beat-for-beat (shock, laugh, deadpan) |
| Sing/perform-along duet | Duet | Original is a song/sound; user harmonizes, dances, or mirrors |
| Versus / side-by-side | Duet | Original vs user — same prompt, two outcomes, comedic contrast |
| Co-sign / amplify duet | Duet | Original makes a point; user nods/gestures agreement, captions add commentary |
For tech/AI niches, stitch architecture is the strongest hook — cold-open on the original's claim, hard cut to the user's reaction. Cross-reference the stitch hook guidance. If a trend doesn't fit an archetype, force it into the closest one.
Stage 0 — Discovery
Ask in a single message:
1. Instagram or TikTok handle (required). One is enough; both is better. Accept @name, name, or full URL. 2. What do you want this content for? (free text — "grow my brand", "personal account", "promote my SaaS", "just for fun") — biases the niche-vs-broad-viral mix. 3. Comfort on camera? (face + voice / face but silent / faceless) — gates archetypes. Faceless excludes real-time-reaction duets where the face is the punchline; "silent" routes to reaction-by-expression + caption commentary rather than spoken responses. 4. Filming constraints? (e.g. "only at home", "no props", "desk setup only") — optional, biases feasibility.
Do NOT ask: trend count (default 10), niche (derive from scrape), aesthetic (derive from scrape), stitch-vs-duet preference (the trend dictates the layout).
Stage 1 — Personality / niche analysis (auto, no extra prompts)
Once you have the handle, scrape immediately — the handle IS consent.
1. Scrape the profile — mcp__plugin_pika_pika__scrape_social on the handle. Prefer digest: true with digest_top_n: 12 for the first profile read so large user-posts payloads do not flood the context; fetch raw posts only for specific media URLs you need to verify. Pull the most recent 12–20 posts only when the compact result is not enough. Capture: captions, hashtags, post types (reel / carousel / static), recurring locations, recurring people, view counts, music choices, whether they already stitch/duet anything.
2. Fallback if scrape_social returns empty / rate-limited — mcp__plugin_pika_pika__capture_website on https://www.tiktok.com/@{handle} or https://www.instagram.com/{handle}/ for at least the bio + grid screenshot. Tell the user the analysis is grid-only.
3. Identity-confirmation gate before profiling — confirm the scrape or screenshot is the intended creator before you synthesize the profile. Cross-check display name, verified badge, follower count, bio, platform, and whether recent posts match the user's expected creator. Try common handle variants first: with/without dots, dotless, underscores removed, and cross-platform Instagram / TikTok / YouTube checks. Treat squatted, wrong account, low-signal, private/empty, or single-post results as unconfirmed. If the default scrape returns only one stale post but the bio/follower/name signal suggests the account may be real, try user-reels/user-posts with digest and the common variants before calling the account dead. When unconfirmed, stop and ask "Is this you?" with the evidence you saw (N followers, verified badge status, display name, bio snippet, platform URL, recent-post summary) and offer the likely variant; do not synthesize the Creator Profile before identity is confirmed. When identity is confirmed, set state.identity_confirmed = true.
4. Synthesize a Creator Profile (internal, then summarized for the user):
- Niche — primary topic cluster
- Voice & tone — 3 adjectives (dry / earnest / chaotic / aspirational / deadpan / playful / hyped / soft / sarcastic / nerdy)
- Aesthetic — color palette, lighting, framing, indoor vs outdoor, selfie vs third-person
- Caption style — short clipped vs long rambly; lowercase-only? emoji-heavy? all-caps for emphasis? — copy this EXACTLY for on-screen captions
- Reaction style — how do they emote/argue/joke in their existing posts? (deadpan stare, big laugh, rapid-fire rant, slow build) — this drives the reaction script
- Recurring motifs — repeating props, locations, catchphrases, sign-offs
- What works for them — flag the 2-3 posts with disproportionate engagement
- Filming environment baseline — what spaces appear in their grid
5. Present the Creator Profile back to the user in ~6 short lines, then say: "Rolling stitch/duet trend research now — flag anything you want me to recalibrate."
Stage 2 — Find viral videos worth reacting to (you do the legwork)
The agent finds the proven viral originals AND fetches them itself — the user never hunts down a link. For every card on the Stage-3 menu you MUST have already captured a concrete, openable, currently-circulating original-video URL with its real play count.
The hunt is for VIRAL ORIGINALS, not for a replicated template. You're looking for videos that are blowing up / widely recognized right now and that hand the user an obvious clever response. Two complementary angles — run both:
1. Reaction magnets (strongest). Videos people are already stitching/dueting — proof they're reaction-bait.
mcp__plugin_pika_pika__scrape_socialtiktok / keyword—"stitch this","duet this","reply","react to this", plus the user's niche words ("AI video","AI art","is this AI"for an AI creator)tiktok / hashtag—#stitch,#duet,#greenscreen,#react, niche tags- When you find a stitch/duet getting traction, trace it back to the ORIGINAL it reacts to and verify the original's play count.
2. Viral originals + current moments. The big videos/claims/clips of the moment that beg for a take.
tiktok / trending-feedwithparams.region(user's target posting region when known, otherwiseUS) andtiktok / popular-hashtags— pull the genuinely high-play videos. Show the region used on the menu; if the user says another region, re-run the trend scrape for that region before they pick.WebSearchfor this week's viral videos / viral moments / controversial clips / "everyone is talking about" — then verify each on-platform for the real view count (a blog mention is a lead, never proof)instagram / reels-search+instagram / hashtagfor the same
Bias toward originals with an obvious angle for THIS user. A dry-deadpan AI creator gets the most mileage reacting to: AI-skeptic takes ("AI will never make real art"), "is this real or AI" guessing clips, viral fashion/tech hot-takes, absurd internet moments she can deadpan at. The original is vibe-agnostic — the user's response is where their voice goes (cross-reference the trend-vs-voice separation rule).
For every candidate, capture:
- What it is — one line, the recognizable handle ("the guy who says AI art has no soul", "the $7 coffee rant")
- The original video — direct TikTok/IG URL, with its real play count (≥500K; millions ideal)
- The exact beat to stitch — which seconds of the original to show before the cut (the claim / the question / the punchline-setup), e.g. "show 0:00–0:05 where he says '___'"
- The response angle — the one-line take the user fires back (this becomes the script in Stage 4)
- Recency — confirm the original is circulating now (posted/resurging in ~last 30–60 days)
- Reaction-proof (if any) — note existing stitch/duet versions + their traction; it's a strong plus, not a hard requirement
Gate each candidate against "What counts as a TREND here": is the ORIGINAL genuinely viral (≥500K, real number in hand)? recognizable? circulating now? is there an obvious clever response? If yes → it's a card. If the original is low-view, or there's no real take to be had, or you can't name it in a sentence → drop it. If you run out of qualifying originals, deliver fewer than 10 and say so — never pad with low-view clips.
For each card, classify the response density:
| Density | Examples | Filming difficulty |
|---|---|---|
| Single-shot reaction | One locked-off take responding to camera | Low — phone on tripod, one take |
| 2-3 shot mini-response | Setup → reveal, or claim → demo | Low-medium |
| Talking response | User speaks a scripted reply (transcribe for captions) | Medium |
| Show-and-tell / prop | User shows or makes something to answer the original | Medium — prop + framing continuity |
| Real-time duet | User reacts beat-for-beat against the playing original (use duet layout) | Medium — timing to the original matters |
Stage 3 — Present the menu (up to 10 viral videos to react to)
Aim for 10, but only as many as genuinely clear the viral bar — a short honest menu beats a padded one. Mix broad + niche. Never all-broad, never all-niche. Cross-reference the broad-versus-niche menu rule.
- ~4 niche-fit originals (in/around the user's world — AI, art, fashion, internet culture — that still clear the view bar)
- ~4 broad-viral originals (universally recognized moments/clips — bigger reach ceiling)
- ~2 wildcards (an unexpected viral original with a surprisingly good angle for them)
Each card is built around ONE viral original to react to — the proof (the original's play count) is visible right on the card, and you preview the response you'd write:
[1] {What the original is — the recognizable handle, e.g. "AI bro: 'real artists will never use AI'"}
🔥 Viral proof: {original's play count, e.g. "4.1M plays"} • posted/resurging {recency} • {"+ N stitch/duet reactions already" if reaction-proven}
▶ Original video: {direct tappable TikTok/IG URL}
Show this beat: {which seconds to play before the cut — e.g. "0:00–0:05, the 'no soul' line"}
Your response (the angle): {one-line preview of the take you'll script in their voice — e.g. "cut to you, deadpan: 'made this in 4 seconds. anyway' + reveal"}
Layout: {Stitch (default) / Duet} • Response density: {single / 2-3 / talking / show-and-tell / real-time}
Requirements before picking: {none OR required prop/disclosure/source constraint/portrait filming note, stated plainly}
Why it's a fit: {1 line — why this original + this angle lands for them}Every card MUST show: what the original is, its real play count (≥500K), a working URL, and the beat to stitch. The proof is the ORIGINAL's virality — no play count, no card. Mix broad-viral (universally recognized originals) with niche-fit (originals in/around the user's world that still clear the view bar). If fewer than 10 reaction-worthy viral originals clear the bar, deliver fewer and say "only N viral originals clear the bar right now" — never pad with low-view clips. Cross-reference the trend fingerprint gate, the virality receipts gate.
Then ask: "Which one do we make? (pick a number)"
Do NOT write the full script or fetch/probe the original yet — wait for the pick. Saves work if they want something else.
Stage 4 — Full production package (per trend chosen)
When the user picks a trend, deliver the package below in one message, then walk them through filming → compositing.
4a. Reaction / continuation script + concept options (in their voice, MULTIPLE variations)
Deliver 6-10 response variations for the user to pick from — never just one. Each variation pairs one reaction/continuation script (what the user says or does) with the concrete visual concept it films against. Write all in the user's voice from Stage 1:
- Mirror their voice, cadence, and reaction style EXACTLY — if they're deadpan, the response is deadpan; if they rant, it rants. Use their catchphrases / sign-offs.
- Match the trend's structure (the stitch sets up, the user pays off; the duet reaction lands on the original's beat) but the content is theirs.
- For stitches: the user's response must make sense cutting from the original's segment. Write the response to land its hook in the first 1-2 seconds after the cut — that's where retention is decided.
- For talking responses: keep spoken length to the trend's response length (most 8–20s). Give a one-line delivery direction ("first sentence fast, then drop tempo").
- For performance duets: note the beat the user mirrors/hits relative to the original's timeline.
- Cross-reference the multi-option script and shot-list contract — multi-option scripts + explicit filming breakdown are MANDATORY for every Stage 4 delivery.
Save the chosen reaction script as state.script_text once the user picks a variation. Keep any broader concept metadata separately; the teleprompter URL reads state.script_text directly.
4b. Filming breakdown — MANDATORY (exact filming directions)
The user MUST know exactly what to shoot. Non-negotiable for every Stage 4 delivery. Numbered shots, each filmable on the user's phone with no crew. For each:
- Shot # and duration (s) — for duets, this must match the original's length (they play simultaneously); for stitches, the user's clip can run as long as the response needs.
- Camera position — tripod / leaned / handheld — and HEIGHT (eye level, chest, low angle).
- Where to stand / sit — distance from camera (feet/cm), orientation (facing lens / 3⁄4 / side).
- Framing — and for duets, frame for the split: the user occupies one panel, so shoot with the subject biased toward the inner edge (toward the original) and leave headroom — a tight-but-not-cramped medium/close works best in a split panel. Tell them which side they'll be on (see Stage 6 for the convention).
- Action / performance — exact micro-movement with timing ("at the cut, look dead into lens, hold 1s, then 'absolutely not'"). For duets, tie reactions to the original's beats ("laugh when the original says 'and then it deployed itself'").
- Props in frame — every prop named and placed.
- Look direction — at lens / off-camera / at the original (for duets, many creators glance toward the original's side as if watching it).
- Lighting — window / desk lamp / ring light; key position relative to face.
- Wardrobe note — one line, only if it matters.
Be filmable, not vibey. Cross-reference the multi-option script and shot-list contract.
4c. Caption layout (timed + positioned)
For each on-screen caption, specify:
- Text (exact words, casing, punctuation/emoji — the user's style).
- In/out timing —
t=0.0s → 2.4s, tied to the stitch cut or the duet beats. - Position — bottom safe zone (y ≈ 1255 for 1080×1920 stitch output) by default. For duets, captions go in the bottom safe zone of the final split-screen output, centered across the composite and clear of the vertical split. Do not assume a fixed x-coordinate; verify the output dimensions with
mcp__plugin_pika_pika__probe_mediabefore captioning. Cross-reference the caption safe-zone guidance for the placement table. - Style —
reels-clean, bold white text, 4px black outline per the caption safe-zone guidance. Deviate only if the trend has a signature caption style (call it out).
4d. The original clip + audio plan
- Restate the original-clip URL (from the menu card) so the user can sanity-check.
- Stitch point (for stitches) — the exact original segment used (e.g. "0:00–0:03"). Save it as
state.original_segment_start_sandstate.original_segment_end_s; every stitch trim, caption offset, and optional post-hook bed slice must use those same bounds. - Audio plan — state which audio survives in the final cut:
- Stitch: original segment keeps its own audio; hard-cut to the user's clip with the user's own audio. (Sequential — no mixing needed unless a continued bed is wanted.)
- Reaction duet: original audio ducked to ~−10 dB; user's voice/audio at 0 dB on top.
- Performance / sing-along duet: original audio full (it's the trending sound); user's mic off or low — the user performs to it.
4e. Filming checklist (handed to the user)
- Phone in airplane mode + Do Not Disturb.
- Vertical / portrait orientation locked.
- 1080p 60fps (or 4K 30fps).
- Clean lens; lock exposure (tap-and-hold) before filming.
- For duets — play the original on a second device while filming so reactions/performance land on the right beats. (You'll still composite from the clean source, but performing to it gets the timing right.)
- For duets — match the original's length (film a take at least as long as the original).
- Film each shot 2-3 times for cutaways.
4f. Teleprompter handoff (hand the reaction script to the user's phone)
Once the user picks a reaction script variant from Stage 4a and approves it — hand them the canonical Pika teleprompter via formats/teleprompter.md so they can record their half with the script scrolling on their phone (under the lens, per-line pacing, 3-2-1 countdown, Upload when done).
For STITCH: the script being prompted is the user's response after the hard cut from the original. Length should at minimum match what they need to say; aim for the same total length as the original they're reacting to.
For DUET (side-by-side): the script is the live reaction running ALONGSIDE the original. They should play the original on a second device for timing, and the teleprompter scrolls their commentary in sync.
Create the handoff through MCP. Do not build a long URL yourself:
handoff = mcp__plugin_pika_pika__create_teleprompter_handoff(
script=state.script_text, # the chosen reaction script variant
handle=state.handle.lstrip("@"),
trend=state.pick.name, # e.g. "@calvin.james41 — everything is AI"
format="duet",
aspect_ratio=getattr(state, "recording_aspect_ratio", "9:16"),
filename=f"{state.handle.lstrip('@')}-duet-take.webm",
mime_type="video/webm", # preferred; upload-return accepts browser mp4/webm variants
max_size_bytes=350_000_000,
expires_in_s=86400,
)
state.teleprompter_url = handoff["teleprompter_url"]
state.teleprompter_qr_image_url = handoff["qr_image_url"]
state.teleprompter_status_url = handoff["status_url"]
state.teleprompter_aspect_ratio = handoff["aspect_ratio"]
url = state.teleprompter_urlDo NOT pass raw presigned media URLs into the teleprompter. Use create_teleprompter_handoff only: it creates the browser-safe upload-return session behind the token, gives the agent a status_url, mints the CDN presign only after the user starts uploading, and avoids TTL failures from recording sessions that take more than a few minutes. The hosted page fetches script, upload_url, and aspect_ratio from MCP, POSTs {mime_type,size_bytes} to upload_url, receives direct_upload_url, attempt_id, and complete_url, PUTs the Blob to the CDN URL, then completes by POSTing attempt_id to complete_url. Keep status_url in agent state only; do not put it in the browser URL.
Emit the URL and returned QR image URL (canonical handoff — see formats/teleprompter.md):
qr_image_url = state.teleprompter_qr_image_url
qr_block = f""Caption to surface to the user (verbatim, swap the original's reference):
📱 Film your reaction on your phone. Scan the QR image or open the link below. Your reaction script is already loaded, with the read zone at the top right under the camera lens. Play the original on a second device for timing.
>
{qr_block}
>
🔗 Or open here: {url}
>
When the take is ready, hit Upload. I'll watch the upload status and compose the stitch/duet as soon as it lands. If upload fails, use Share/Save and send the MP4 back here.
Do not generate a local QR PNG and do not call a third-party QR service; use the qr_image_url returned by create_teleprompter_handoff.
After this — poll state.teleprompter_status_url until it returns status="uploaded" with public_url, then save public_url as both state.user_take_url and state.user_clip_url. state.user_clip_url is the field Stage 5 probes and Stage 6 trims/reframes, so do not leave it unset after a successful teleprompter upload. If upload fails and the user sends a manual attachment, Stage 5's "Receive + sanity-check the user's footage" begins when the file lands.
Stage 5 — Acquire the original + receive the user's footage
This playbook needs two video inputs: the original clip (agent fetches) and the user's response (user uploads).
Fetch the original clip yourself (don't make the user supply it)
- For TikTok/IG public URLs — call
mcp__plugin_pika_pika__scrape_social(tiktok / videoorinstagram / post) withrehost: true; save the returned durable video URL asstate.original_video_urland use that URL directly in the edit tools. - If scrape returns no media URL (private / age-gated / region-locked), tell the user and ask them to upload a local copy of the original.
- Probe the original —
mcp__plugin_pika_pika__probe_media: capture exact duration, resolution, aspect ratio, fps. For stitches, confirm the stitch-point segment exists. For duets, this duration is the target length for the user's clip.
Receive + sanity-check the user's footage
When the user's clip arrives: 1. Probe — save the uploaded clip URL as state.user_clip_url, then call mcp__plugin_pika_pika__probe_media: confirm portrait orientation, ≥1080×1920, duration, fps. The teleprompter should upload a 9:16 take even when the raw camera stream is 16:9; if the clip is still landscape, ask for a phone/portrait reshoot unless the user explicitly accepts a center-crop rescue. 2. Sanity-check against the shot list — was the response captured? For duets, is it ≥ the original's length? If a clip is unusable (wrong orientation, blurry, missing the reaction, too short for a duet), call it out and ask for a reshoot of only that shot. 3. Phone-cameo gate — if the user's clip shows a phone (common in reaction content), confirm it's a current-gen iPhone (15 Pro / 16 / Air). See the phone-cameo gate.
Stage 6 — Compose the stitch / duet + edit
The composite is mechanical once both clips are in hand. Use MCP edit tools for trimming, reframing, layout, audio balance, and caption burn.
6a. STITCH composite (sequential)
1. Trim the original to the stitch segment with mcp__plugin_pika_pika__edit_trim(video_url=state.original_video_url, start_s=state.original_segment_start_s, end_s=state.original_segment_end_s), then save the returned URL as state.original_segment_url and set state.original_segment_duration_s = state.original_segment_end_s - state.original_segment_start_s. Do not use a separate hard-coded hook range here; the stitch video, caption offsets, and optional bed continuation must all share the exact Stage 4d bounds. 2. Normalize both stitch inputs to 9:16 — call mcp__plugin_pika_pika__edit_reframe(video_url=state.original_segment_url, target_aspect="9:16", fill_mode="crop") and save the returned URL as state.original_segment_normalized_url; call mcp__plugin_pika_pika__edit_reframe(video_url=state.user_clip_url, target_aspect="9:16", fill_mode="crop") and save the returned URL as state.user_response_normalized_url. Use fill_mode="pad" only when preserving the full source frame matters more than filling the screen. 3. Concat original segment → user response with mcp__plugin_pika_pika__edit_concat(video_urls=[state.original_segment_normalized_url, state.user_response_normalized_url]), save the returned URL as state.stitch_concat_url, and set state.caption_input_url=state.stitch_concat_url unless optional bed continuation replaces it. Each segment keeps its own audio: the original hook first, then the user's response. 4. Optional bed continuation — only if the original's sound should continue under the response. First call mcp__plugin_pika_pika__extract_audio_from_video on state.original_video_url, then use the probed original duration as state.original_audio_duration_s. Reuse the already-set state.original_segment_start_s, state.original_segment_end_s, and state.original_segment_duration_s from the stitch trim step. Call mcp__plugin_pika_pika__edit_audio_trim(audio_url=<original_audio_url>, start_s=state.original_segment_end_s, end_s=<original_audio_duration_s>) to create state.post_hook_bed_audio_url; end_s is required by the tool schema, so do not omit it. Mix that trimmed bed onto the concat with mcp__plugin_pika_pika__edit_audio_mix(video_url=state.stitch_concat_url, audio_url=state.post_hook_bed_audio_url, original_gain_db=0, audio_gain_db=-15, audio_offset_s=state.original_segment_duration_s), then save the returned URL as state.stitch_mixed_url and set state.caption_input_url=state.stitch_mixed_url. Use -15 dB as the default bed level; adjust within -12 to -18 dB only by passing a single concrete number. Both steps are load-bearing: trimming starts after the exact original segment end in the full source, and the positive offset delays the post-hook bed until the user's response starts in the stitched concat. If edit_audio_trim fails, skip the bed, keep the clean stitch, and set state.caption_input_url=state.stitch_concat_url rather than replaying the hook.
6b. DUET composite (side-by-side, simultaneous)
Convention: original on the right, user on the left (TikTok's duet layout puts your new camera on the left). State this to the user; it's a one-line swap if they want it flipped.
1. Match lengths — both halves play together. Use the Stage 5 probes to set state.original_duration_s and state.user_clip_duration_s, then set state.matched_duration_s=state.original_duration_s unless the user explicitly approved a shorter original segment. If state.user_clip_duration_s < state.matched_duration_s, ask for a longer take instead of freezing a panel. Trim both sources with explicit URLs: call mcp__plugin_pika_pika__edit_trim(video_url=state.original_video_url, start_s=0, end_s=state.matched_duration_s) and save the returned URL as state.original_duet_trimmed_url; call mcp__plugin_pika_pika__edit_trim(video_url=state.user_clip_url, start_s=0, end_s=state.matched_duration_s) and save the returned URL as state.user_duet_trimmed_url. 2. Probe and pre-normalize inputs — call mcp__plugin_pika_pika__probe_media on state.user_duet_trimmed_url and state.original_duet_trimmed_url, then reframe only when a source is not already clean portrait: call mcp__plugin_pika_pika__edit_reframe(video_url=state.user_duet_trimmed_url, target_aspect="9:16", fill_mode="crop") for the user panel and mcp__plugin_pika_pika__edit_reframe(video_url=state.original_duet_trimmed_url, target_aspect="9:16", fill_mode="crop") for the original panel. Save the exact panel-source URLs as state.user_panel_url and state.original_panel_url: use the reframe result when reframe ran, otherwise use the corresponding trimmed URL. These same panel-source URLs must be used for both state.duet_html and the later audio extraction so the rendered panels and mixed tracks come from the same timeline. 3. Render the native 9:16 side-by-side visual — author state.duet_html as a HyperFrames HTML document and call mcp__plugin_pika_pika__render_html_animation(html=state.duet_html, format="mp4", fps=30, quality="standard"). The root composition must be data-composition-id="content-director-duet", data-width="1080", data-height="1920", and data-duration="<matched_duration_s>". Put the user's video in a left panel at x=0,width=540,height=1920 and the original in a right panel at x=540,width=540,height=1920, with each video using cover-crop/object-fit semantics so neither panel stretches. Include a timed class="clip" composition child with data-start="0", data-duration="<matched_duration_s>", and stable data-track-index, plus a real window.__hf.seek(t) implementation that seeks both video elements before the renderer captures the frame. Minimal template:
<div data-composition-id="content-director-duet" data-width="1080" data-height="1920" data-duration="<matched_duration_s>">
<div class="clip" data-start="0" data-duration="<matched_duration_s>" data-track-index="0">
<video id="user-panel" src="<state.user_panel_url>" muted playsinline preload="auto" style="position:absolute;left:0;top:0;width:540px;height:1920px;object-fit:cover;"></video>
<video id="original-panel" src="<state.original_panel_url>" muted playsinline preload="auto" style="position:absolute;left:540px;top:0;width:540px;height:1920px;object-fit:cover;"></video>
</div>
</div>
<script>
const duration = Number("<matched_duration_s>");
const videos = Array.from(document.querySelectorAll("video"));
const ready = Promise.all(videos.map((video) => new Promise((resolve) => {
if (video.readyState >= 2) return resolve();
video.addEventListener("loadeddata", resolve, { once: true });
})));
async function seek(t) {
await ready;
const target = Math.max(0, Math.min(Number(t), duration));
await Promise.all(videos.map((video) => new Promise((resolve) => {
video.pause();
if (Math.abs(video.currentTime - target) < 0.001 && video.readyState >= 2) return resolve();
video.addEventListener("seeked", resolve, { once: true });
video.currentTime = target;
})));
}
window.__hf = { duration, ready, seek };
</script>Save the returned URL as state.duet_visual_url. 4. Add the duet audio with `mcp__plugin_pika_pika__edit_audio_mix` — extract audio from state.user_panel_url and state.original_panel_url with mcp__plugin_pika_pika__extract_audio_from_video, then mix tracks onto state.duet_visual_url with original_gain_db=-60 so any audio captured by the HTML render cannot double the extracted mix. Save the returned URL as state.duet_mixed_url and set state.caption_input_url=state.duet_mixed_url:
- Reaction duet:
tracks=[{url: user_audio_url, gain_db: 0}, {url: original_audio_url, gain_db: -10}]. - Performance / sing-along duet:
tracks=[{url: original_audio_url, gain_db: 0}], adding{url: user_audio_url, gain_db: -18}only if the user's mic should remain faintly audible. - No dialogue / silent reaction: use the original audio track unless the user's on-set sound is explicitly part of the bit.
MCP caveat: mcp__plugin_pika_pika__edit_split_screen is not the native duet production path because horizontal split-screen is a co-equal stacked output shape rather than a guaranteed 1080×1920 half-canvas render. Use it only if the user explicitly accepts that fallback, and always probe/disclose the returned dimensions. Do not crop a finished split-screen just to force 9:16; that can cut off one panel.
6c. Burn captions with mcp__plugin_pika_pika__add_captions
Use mcp__plugin_pika_pika__add_captions on state.caption_input_url from 6a/6b, never on the pre-audio visual render or pre-mix stitch concat.
- If the user speaks a scripted response (talking stitch / reaction with dialogue): transcribe the same user timeline that will appear in the final composite. For stitches, transcribe
state.user_response_normalized_url; offset those caption rows bystate.original_segment_duration_sso timings line up after the original hook. For side-by-side duets, transcribestate.user_panel_urlorstate.user_duet_trimmed_url, never the longer raw upload; drop any row whose start timestamp is at or beyondstate.matched_duration_s, and clamp each remaining row's end timestamp tostate.matched_duration_s. Buildstate.caption_rowsas[{start_s, end_s, text}]segment/phrase rows, then callmcp__plugin_pika_pika__add_captions(video_url=state.caption_input_url, caption_mode="manual", subtitles=state.caption_rows, style="reels-clean", position="bottom", margin_v=665, font_color="white", outline_color="black", outline_width=4). - If the response is silent (expression-only reaction, performance): pass the title-card caption rows from the 4c layout in manual mode.
After the final caption pass, read the returned url and save it as state.final_url. Do not return state.duet_visual_url, the pre-caption audio mix, or any intermediate stitch concat as the deliverable.
Caption params (validated):
style="reels-clean"font_size≈57(default; 50-72 per readability)font_color="white",outline_color="black",outline_width=4position="bottom",margin_v≈665for the bottom safe-zone default- For duets, keep captions centered across the final composite so they clear the vertical split; use
mcp__plugin_pika_pika__probe_mediaafterrender_html_animationbefore picking margins. - Keep caption rows short enough to render without clipping. Split long rows manually before
add_captions; do not rely on automatic wrapping for long one-line punchlines.
IG Reels safe zone (1080×1920 stitch/native duet output): top unsafe y≈0-270, bottom unsafe y≈1480-1920. Captions sit inside y≈270-1475, ~80px margin from left/right edges. Default bottom visual placement is near y≈1255. For explicitly accepted edit_split_screen fallback output, scale the same safe-zone intent to the probed output height. Emoji warning: decorative emoji can render inconsistently in burned captions; put emoji in the post description at upload.
Pre-delivery gates
1. Composite-integrity gate — open the final. For stitches: the cut is clean, the original segment is the right hook, audio doesn't clip at the join. For duets: the split is even inside the 1080×1920 render, neither panel is stretched, both play in sync, output dimensions match mcp__plugin_pika_pika__probe_media, and audio mix is balanced (you can hear the user over the original in a reaction duet; you hear the original in a performance duet).
2. Caption legibility gate — readable at phone-screen size on first pass; doesn't overlap a face or the mouth/chin; for duets, centered and clear of the split. No row is clipped by the left/right edge. Split long captions across two cards, reduce font size within 50-72, or adjust safe-zone margin and re-render. See the caption safe-zone guidance.
3. Audio-sync gate — for stitches, does the user's hook land right after the cut? For duets, do the user's reactions land on the original's beats? Re-trim if not.
4. Orientation gate — verify final portrait output with mcp__plugin_pika_pika__probe_media. If a source clip is rotated or framed wrong, re-run mcp__plugin_pika_pika__edit_reframe before compositing.
5. Phone-cameo gate — any phone visible in the user's half must be a current-gen iPhone (15 Pro / 16 / Air). See the phone-cameo gate.
Output: for stitches and native side-by-side duets, 1080×1920 H.264/AAC, 9:16. If the user explicitly accepted the edit_split_screen fallback, record the probed dimensions from mcp__plugin_pika_pika__probe_media and disclose that it is not the native half-canvas deliverable. Return state.final_url with version label {user_slug}_{trend_slug}_v{N} — never overwrite. Plus a post caption + hashtag set in the user's voice (text-only, lives in the description).
Stage 7 — Loop
After delivering one stitch/duet, ask: "Want to do the next one? (pick another number from the menu, or 'new trends' to re-research)".
If they pick another, return to Stage 4 with that trend. Do NOT re-run trend research unless they ask — the Stage-3 menu of 10 stays warm for the session.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
mcp__plugin_pika_pika__scrape_social → provider_unavailable / rate-limited | Upstream scraper outage | Retry after 30–60s; if it persists, switch the discovery angle (trending-feed ↔ keyword ↔ hashtag) and tell the user research is degraded — don't fabricate cards to fill the gap |
| Scrape returns no media URL for the chosen original | Private / age-gated / region-locked source video | Tell the user and ask them to upload a local copy of the original |
mcp__plugin_pika_pika__transcribe_audio returns one big segment | ASR emitted coarse segment timing | Use the approved reaction script to create manual phrase rows, align them to the planned delivery beats, and keep the uncertainty visible in the delivery note |
| User clip plays rotated / sideways after composite | iPhone rotation metadata was not normalized before compositing | Verify with mcp__plugin_pika_pika__probe_media; if pixels are still rotated, re-run mcp__plugin_pika_pika__edit_reframe before composing |
mcp__plugin_pika_pika__edit_concat errors or audio drops at the join | Original & user clips have incompatible audio layouts | Re-run mcp__plugin_pika_pika__edit_reframe on both sources, then retry mcp__plugin_pika_pika__edit_concat; if the original has no audio, keep the hook short and surface that caveat |
| Original already has burned captions + narrator audio (AI brainrot, subtitled clips) | Source isn't clean | Keep the stitch beat short (≤ the hook) so it reads as "the chaos"; don't fight it — your captions only go on the user's half |
| Menu comes back under 10 cards | Few originals clear the ≥500K viral bar this week | Deliver fewer and say "only N viral originals clear the bar right now" — never pad with low-view clips |
| Original is taller/shorter than 9:16 (e.g. 576×1048, 1:1, landscape) | Source aspect ≠ 1080×1920 | Run mcp__plugin_pika_pika__edit_reframe(target_aspect="9:16", fill_mode="crop"); never stretch |
| User teleprompter take is landscape | User opened the link on a desktop camera or the browser ignored portrait capture | Ask for a phone/portrait reshoot unless the user accepts center-crop rescue; do not treat landscape as the expected path. |
| Caption row clips at the frame edge | Long punchline sent as one row | Split the caption into shorter timed rows before re-running mcp__plugin_pika_pika__add_captions. |
Don'ts
- Don't propose formats with no original clip to pair against. This playbook is stitch/duet only — the user's footage MUST combine with an existing creator's video. Plain talking-head →
formats/talking.md. Silent POV →formats/pov.md. AI dance →formats/dance.md. Multi-format → the Content Director front door. - Don't make the user find the original. The agent finds AND fetches the viral original itself (Stage 5) — the user only films their response.
- Don't invent originals. Every card needs a real, tappable original-video URL with a real play count in hand. If research comes back empty, tell the user and broaden — don't fabricate.
- Don't require a replicated template. This is the big one: a stitch/duet reaction is NOT a fingerprinted trend that 10 people copy identically — everyone reacts to the SAME viral original in their OWN way. The proof is the ORIGINAL's virality, not a replication wave. Don't drop a great reaction-worthy viral video just because nobody's stitched it in a specific repeated format.
- Don't pitch a low-view original. Hard floor on the ORIGINAL: ≥500K plays (millions ideal). The entire point is borrowing a video the algorithm and audience already recognize. A low-view source = nothing to borrow. Drop it.
- Don't pitch a vague topic instead of a specific video. "React to AI drama" is not a card. "React to {this specific 4M-play clip}, show 0:00–0:05, here's your line" is a card. Always a specific, nameable, viral source video.
- Don't pad the menu to hit 10. If fewer than 10 viral reaction-worthy originals clear the bar, deliver fewer and say "only N clear the bar right now." Never inflate with low-view clips or blog-citation-only entries.
- Don't pitch a stale moment. The original must be circulating now (last ~30–60 days) or in an active resurgence. A viral moment whose wave passed months ago is dead air.
- Don't strip the user's voice from the response. The original is the original; the response is 100% theirs — casing, punctuation, catchphrases, deadpan/rant cadence. Cross-reference the trend-vs-voice separation rule.
- Don't stretch the clips in the composite. Always cover-crop (or letterbox) before composing — never distort aspect to fill a panel. For native side-by-side duets, render the 1080×1920 half-canvas via
mcp__plugin_pika_pika__render_html_animation; useedit_split_screenonly as a disclosed fallback. - Don't desync a duet. Both halves must be the same length and play together. Equalize duration before
mcp__plugin_pika_pika__render_html_animation; trim, don't let one half freeze mid-reaction. - Don't bury the user under the original in a reaction duet. The user's audio sits ON TOP (0 dB), the original ducks (~−10 dB). For performance duets it's the reverse. Pick per the 4d plan.
- Don't burn captions outside the safe zone, and for duets keep them centered and clear of the vertical split. For 9:16 stitch/native duet output, stay inside y=270-1475. For explicitly accepted split-screen fallback output, scale the same safe-zone intent to the probed dimensions. See the caption safe-zone guidance.
- Don't chain arbitrary text overlays for captions. Use one
mcp__plugin_pika_pika__add_captionspass withreels-clean, 4px black outline, and safe-zone placement — see Stage 6c and the caption safe-zone guidance. - Don't claim the composite gets native TikTok stitch/duet attribution. It doesn't — be honest (see the native-feature reality section). It's the only path on Reels and a controllable path on TikTok; the in-app button is the native-credit alternative.
- Don't show old phones. Any phone in the user's footage must be a current-gen iPhone (15 Pro / 16 / Air). See the phone-cameo gate.
- Don't skip the Stage 3 menu and jump to production. Present the menu (up to 10), always wait for a pick. Producing without a selection burns the user's budget and tokens on a video they may not want.
Content Director — Teleprompter (browser camera + prompter add-on)
What this does
The formats/talking.md and formats/duet.md playbooks both reach a point where they hand the user a script and say "now film it." That hand-off used to be a dead-end — the user had to copy the script into a phone teleprompter app, set up framing, fight with prompter speed, and remember to send the result back.
The handoff replaces that step with a one-tap short link. The agent calls mcp__plugin_pika_pika__create_teleprompter_handoff and sends the returned URL:
https://teleprompter.pika.bot/r?t=...The user taps it → camera + prompter come up → they hit record → re-shoot → Upload. The page uses the token to fetch script, upload_url, and aspect_ratio from MCP. It POSTs {mime_type,size_bytes} to that browser-safe upload_url. The start response returns direct_upload_url, direct_upload_headers, attempt_id, and complete_url; the page PUTs the recorded Blob directly to the CDN URL, then POSTs the attempt_id to complete_url. The agent polls the returned status_url until it sees public_url, then resumes its pipeline (captions, edit, etc.). If upload fails, the page falls back to Share/Save so the user can still send the take manually.
Upload-return flow
The canonical playbook handoff uses mcp__plugin_pika_pika__create_teleprompter_handoff. Do NOT pass raw `upload_asset` presigned URLs to the browser: short-lived presigned URLs can expire during a realistic recording session (open link → adjust speed → 2-3 takes → preview → upload), and they put a storage bearer URL in front of the user before the user is ready to upload.
So for this flow: call create_teleprompter_handoff, emit its short teleprompter_url, and keep its status_url in agent state only. The MCP handoff row stores the script, upload_url, and aspect_ratio; the browser fetches those through /api/teleprompter-handoff?t=.... Do not pass the agent-pollable status_url to the browser. The page sends a JSON start request to upload_url, receives direct_upload_url, direct_upload_headers, attempt_id, and complete_url, PUTs the Blob to the CDN URL, then completes the session by POSTing {attempt_id}. The MCP upload-return endpoint mints the short-lived CDN presign only when upload starts, so recording time no longer races the presign TTL.
Used by
- ✅
formats/talking.md(Stage 4f) — script is the spoken delivery to camera - ✅
formats/duet.md(Stage 4f) — script is the user's reaction half of the stitch/duet - ❌
formats/pov.md— silent acting, deliverable is on-screen captions + shot list - ❌
formats/dance.md— AI-generated dance from a photo, no live filming
The short-token contract
create_teleprompter_handoff returns:
| Field | Who uses it | What it does |
|---|---|---|
teleprompter_url | User/browser | Short URL, normally https://teleprompter.pika.bot/r?t=.... This is what the agent surfaces to the user and QR generator. |
qr_image_url | User/agent | Browser-viewable SVG QR image for the same teleprompter_url. Render this in chat so desktop users can scan with a phone. |
status_url | Agent only | Poll until it returns status="uploaded" with public_url. Never show this URL to the browser. |
aspect_ratio | Agent + browser | Recording ratio used by the page and shown to the user. Supported: 9:16, 16:9, 1:1, 4:5. Default: 9:16. |
expires_at | Agent | User-facing expiry if needed. |
The browser then reads /api/teleprompter-handoff?t=..., which returns only browser-safe state: script, optional handle/trend/format, aspect_ratio, upload_url, and expiry. It does not return status_url.
Legacy/manual fallback: the page still accepts script/s, handle, trend, format, aspect_ratio/aspect, and #upload=... for local testing. Production playbooks should use the short-token MCP tool so QR codes stay sparse.
Wiring into the playbooks
In formats/talking.md or formats/duet.md Stage 4f (or wherever the script is finalized), add a step:
1. Create the teleprompter handoff:
handoff = mcp__plugin_pika_pika__create_teleprompter_handoff(
script=state.script_text,
handle=state.handle,
trend=state.pick.name,
format=state.format, # "talking" or "duet"
aspect_ratio=getattr(state, "recording_aspect_ratio", "9:16"),
filename=f"{state.handle.lstrip('@')}-{state.format}-take.webm",
mime_type="video/webm", # preferred; endpoint accepts browser mp4/webm variants in the same family
max_size_bytes=350_000_000, # enough for a short 9:16 phone take
expires_in_s=86400,
)
state.teleprompter_url = handoff["teleprompter_url"]
state.teleprompter_qr_image_url = handoff["qr_image_url"]
state.teleprompter_status_url = handoff["status_url"]
state.teleprompter_aspect_ratio = handoff["aspect_ratio"]2. Emit `state.teleprompter_url` and `state.teleprompter_qr_image_url` to the chat — see the next section. 3. Tell the user to record, hit Upload, and return to chat. If upload fails, the page opens the Share/Save fallback. 4. Poll state.teleprompter_status_url until it returns status="uploaded" with public_url, then resume the pipeline with that uploaded take.
The page itself is stateless — no DB, no auth, no analytics. The temporary handoff state lives in the MCP jobs registry behind the short token.
The handoff — URL + returned QR image (canonical pattern)
Every time the agent hands off the link, show both values returned by create_teleprompter_handoff:
1. The clickable URL — for desktop users who want to record at their workstation, or who'd rather paste the link into their phone manually. 2. The `qr_image_url` SVG — for desktop users who want to record on their phone. Scanning it opens the same URL with the same script preloaded. No copy-paste, no typing, and no local Python package dependency.
Do not invent a local helper name and do not hard-code an unapproved third-party service in the skill:
qr_image_url = state.teleprompter_qr_image_urlThen render it in the handoff message:
Caption the QR with the same message as the URL — e.g. "Film it on your phone — scan or click." If the QR image fails to render in a client, keep the clickable URL visible as the fallback.
URL length budget
QR codes get progressively denser as URLs grow. The short-token URL is normally far below the risky range, which is why this handoff exists. Practical thresholds:
| URL length | Result |
|---|---|
| < 600 chars | Sparse QR, scans instantly even in dim light |
| 600 – 1200 | Dense, still reliable |
| 1200 – 1800 | Very dense — needs a clean phone screen and steady hand. Last reasonable tier. |
| > 1800 | Skip the QR, surface only the clickable URL with a "tap to open on this device" instruction |
If the URL is going to exceed 1500 chars, something is wrong: do not put the script or upload URL back into the QR. Use create_teleprompter_handoff so the script stays server-side behind the token.
Example handoff message
🎬 Your script is ready. Film it on your phone:
📱 Scan this QR on your phone (opens teleprompter.pika.bot with the script preloaded):

🔗 Or open directly:
{state.teleprompter_url}
When you're done, hit Upload. I'll watch the upload status and start the edit as soon as the take lands. If upload fails, use Share/Save and send the MP4 back here.Hosting
Two clean options:
- `web_publish` — push
teleprompter.htmlto Pika's static host and use the returned HTTPS
URL forever. Best for production.
- `capture_website` (dev only) — not useful here; this page needs real camera access from the
user's device.
For local development: python3 -m http.server 8443 --bind 127.0.0.1 won't work for mobile testing because mobile browsers refuse camera on non-HTTPS, non-localhost origins. Easiest mobile dev loop is web_publish to a temporary URL.
What the page does on its own
- Asks for camera + mic on tap (iOS Safari requires a user gesture for
getUserMedia). - Renders the script with each line broken on natural pauses (
. , ! ? —), then force-splits any
chunk over 7 words at the most natural break-word (and / or / to / per / of / from / etc.) near the middle so 20-word run-ons don't survive. Tiny orphan fragments (under 8 chars or a leading conjunction/article on its own) get merged back into the previous chunk so you don't get "yeah," sitting alone.
- Anchors a glowing read zone at the very top of the prompter window — close to the front
camera. The line crossing that zone grows + brightens; lines that have already passed fade and shrink. The user's eye-line stays high.
- Lead-in: the first line is parked ~1.5 line heights below the read zone so when scrolling
starts after the countdown, line 1 travels visibly upward into position — giving the user a beat to lock onto it before they have to speak. (Without this it lands on the read zone the instant REC fires, which is too sudden to track.)
- Per-line pacing: scroll velocity is computed per line as `(distance to next line) ÷
(this line's word count ÷ wpm × 60)`, floored at 0.55s, so a 1-word zinger like "Cents." gets time to breathe and a 10-word sentence scrolls past at the right reading speed. The same WPM setting means very different pixel speeds in different parts of the script — by design.
- Picks the best supported MIME (
video/mp4;codecs=h264,aacfirst on iOS 17+, falls back to webm
on Android/desktop).
- Shows the requested recording ratio on the start screen. Default is
9:16(1080×1920), but the
handoff can request 16:9 (1920×1080), 1:1 (1080×1080), or 4:5 (1080×1350). The page records through that target-aspect canvas with cover-crop normalization, so a mismatched physical camera stream still uploads the requested take shape.
- Records via
MediaRecorder, plays back the take, and lets the user re-shoot or upload. - If the handoff provides
upload_url, starts the upload-return session with JSON, PUTs the Blob to
the returned direct_upload_url, POSTs attempt_id to complete_url, and falls back to Share/Save if any step fails.
Controls (built in)
All controls live on the main record screen — nothing important is hidden in a menu.
| Control | What |
|---|---|
| ● REC | Big red button. Tap → 3-2-1 countdown → recording starts; prompter starts scrolling. Tap again → stop → preview screen. |
| 00:00 timer | Sits just above the slider row when recording — never overlays the script. |
| SPEED slider (60–240 wpm) | Live on the main screen. Adjustable mid-take without opening any menu. |
| SIZE slider (22–46 px) | Live on the main screen. Adjustable mid-take. |
| ⟲ Re-shoot | Discards take, re-parks the first line below the read zone, ready to go again. |
| Upload / Share / Save | Preview action. With upload-return it uploads first; without upload-return or after upload failure it uses native Share or download. |
| Tap the prompter | Pause / resume scroll during recording. In Manual mode (settings), tapping advances one line. |
| ⇋ Mirror | Top-right toggle and also in settings — preview only; the recorded video is canonical. |
| ⚙ Settings sheet | Slides up from the bottom — only contains the two toggles (Mirror, Manual). The sliders are NOT here. |
| ✕ | Top-left bail-out — confirms before discarding. |
What this playbook is NOT
- Not a backend. No accounts, no recording-history, no streaming. One-page, stateless.
- Not a production switcher. It's a single-person front-camera setup.
- Not a beam-splitter prompter. The read zone is the best compromise — eye-line drift is minimal
but not zero.
Triggers (in case the user invokes it directly)
Mostly you'll never trigger this page standalone — the parent playbook calls it. But if the user says any of:
- "give me a teleprompter for this script"
- "open the camera + prompter"
- "set up the recording page"
- "let me film this in the browser"
…produce a URL with at least script= and hand them the link.
Content Director — Bundle
A complete trend-research -> script -> camera -> edit pipeline for short-form video, packaged as one registered Content Director skill with nested format playbooks. The user gives you their IG or TikTok handle. You discover real, currently-viral trends across four formats, hand them the production package, route them into the recording flow, and ship the final edited mp4.
What's in the bundle
content-director/ ← top-level routing skill (this is where it starts)
├── SKILL.md ← registered `/pika:content-director` router
├── formats/
│ ├── talking.md ← talking-to-camera (spoken to lens)
│ ├── pov.md ← silent POV (acting + on-screen captions)
│ ├── duet.md ← stitch / duet (react to a viral original)
│ ├── dance.md ← AI-generated dance from one photo
│ ├── teleprompter.md ← teleprompter handoff contract
│ └── teleprompter.html ← browser camera + scrolling prompter source
└── README.md ← this fileOnly content-director/SKILL.md is registered as a user-invoked skill. The format files under formats/ are playbooks the router loads after Stage 0 locks the format. The teleprompter HTML source is the same static page deployed to https://teleprompter.pika.bot/.
How the user reaches each skill
| User intent | Entry point | Trigger phrases |
|---|---|---|
| "Be my content director" / "what kind of trend should I make" | content-director | The orchestrator; Stage 0 figures out which of the four formats fits, then routes. |
| Already knows they want a spoken hot-take / storytime / hook | content-director talking | Loads formats/talking.md. |
| Wants silent POV acting | content-director pov | Loads formats/pov.md. |
| Wants to react to a specific viral video | content-director duet | Loads formats/duet.md. |
| Wants AI dance from one photo | content-director dance | Loads formats/dance.md. |
The teleprompter is not user-invokable. It's a sub-tool the talking and duet skills hand off into during Stage 4f (after script approval). Its contract is documented in formats/teleprompter.md; the deployed page source is formats/teleprompter.html.
End-to-end flow
┌───────────────────────────────────────┐
│ content-director (Stage 0-1) │
│ - intake handle │
│ - recommend OR show cross-format │
│ sampler │
│ - route │
└──────────────┬────────────────────────┘
│
┌───────────────────┬────────────────┴───────────────┬────────────────────┐
▼ ▼ ▼ ▼
talking pov duet dance
Stages 0-7: Stages 0-7: Stages 0-7: Stages 0-7:
profile, profile, profile, profile,
trend research, POV trend research, find viral original, dance trend research,
menu, script menu, captions+shot list, menu, reaction menu, motion-control,
script,
│ │ │ │
▼ ▼ ▼ ▼
[Stage 4f] (no prompter — [Stage 4f] (no prompter —
TELEPROMPTER shot list TELEPROMPTER AI-generated,
HANDOFF handed directly) HANDOFF no filming)
│ │
▼ ▼
user records on phone (teleprompter.pika.bot/r?t=...)
hits Upload → MCP upload-return → agent polls status_url
│ │
└─────────────────┬──────────────────────────────┘
▼
Stage 5: receive take
Stage 6: edit + captions + audio mix
Stage 7: deliver + loopThe teleprompter — what makes it special
A single static HTML page at teleprompter.pika.bot, backed by a short-token MCP handoff. The agent stores the approved script, browser upload_url, and aspect_ratio in the MCP jobs registry, then gives the user a sparse QR-friendly URL like https://teleprompter.pika.bot/r?t=... plus a returned qr_image_url for phone scanning.
Key behaviors:
- Top-anchored read zone — the line you're reading sits at the very top of the screen, right
under the front-facing camera lens. The line crossing the read zone grows + brightens; lines above fade and shrink. The user's eye-line stays high, not down at their lap.
- Per-line pacing — scroll velocity is computed per chunk as `(distance to next line) ÷
(this line's word count ÷ wpm × 60)`. A 1-word zinger like "Cents." dwells 0.55s; a 10-word sentence dwells ~4.3s at 140 wpm. Same WPM, very different pixel speeds — by design.
- 1.5-line lead-in — the first line parks below the read zone so when scrolling starts after
the 3-2-1 countdown, line 1 visibly travels upward into position. No "blink and miss" start.
- Smart line breaking — splits on
. , ! ? —and force-breaks chunks over 7 words at the
most natural break word (and, or, to, per, of, etc.) near the middle. Orphans (yeah, alone) get merged back.
- Mirrored output — when the mirror toggle is on, recording is routed through a hidden
<canvas> with ctx.scale(-1, 1), so the recorded MP4 matches what the user saw in preview.
- Upload-return first, Share fallback — the page uploads the take through the browser-safe
upload_url; if that fails, navigator.share({ files: [...] }) opens the iOS share sheet (AirDrop, Save Video to Photos, Messages, Mail) on iOS Safari 14+ and falls back to a download on desktop browsers.
- Protocol-selected recording ratio — the handoff can request
9:16,16:9,1:1, or4:5.
The page shows that ratio before recording and records through a matching canvas.
- Live controls — SPEED and SIZE sliders always visible above the record button (stacked on
phone, side-by-side on tablet+). No hunting in menus mid-take.
The deployed URL is production-stable. Each session changes only the token; the underlying app is the same. The token expires with the MCP handoff/upload-return session.
Updating the teleprompter
Source lives in formats/teleprompter.html. To redeploy:
1. mcp__plugin_pika_pika__upload_asset with filename: teleprompter.pdf, mime_type: application/pdf (HTML isn't allow-listed by upload_asset; PDF mime is accepted and the bytes are served as text/html when web_publish republishes them at index.html). 2. PUT the HTML bytes to the returned presigned_url. 3. mcp__plugin_pika_pika__web_publish with slug: pika-teleprompter, files: [{path: "index.html", source_url: <public_url from step 1>}]. 4. Verify with curl -sI https://teleprompter.pika.bot/ — content-type should be text/html.
If you're picking this up as a maintainer, start with content-director/SKILL.md to understand the routing model. The two heaviest playbooks are formats/talking.md and formats/duet.md; they're each self-contained pipelines with stage-by-stage guidance.
Related skills
FAQ
Which video formats does content-director cover?
Four: talking-to-camera, silent POV, dance, and stitch/duet. It does not cover carousels or transitions.
What does content-director need to start?
An Instagram or TikTok handle, and it uses Pika MCP tools for trend scraping, media editing, and the teleprompter handoff.