
Ce Plan
- 1 installs
- 23.9k repo stars
- Updated August 5, 2026
- everyinc/compounding-engineering-plugin
This is a copy of ce-plan by everyinc - installs and ranking accrue to the original listing.
Helps with productivity & planning tasks.
About
ce-plan is a Claude Code skill for productivity & planning. It helps you ship faster with AI-assisted development.
- ce-plan
- Productivity & Planning
- AI-coding skill
Ce Plan by the numbers
- 1 all-time installs (skills.sh)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/everyinc/compounding-engineering-plugin --skill ce-planAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 23.9k |
| Last updated | August 5, 2026 |
| Repository | everyinc/compounding-engineering-plugin ↗ |
What it does
Helps with productivity & planning tasks.
Files
Create Technical Plan
Note: The current year is 2026. Use this when dating plans and searching for recent documentation.
ce-brainstorm defines WHAT to build. ce-plan defines HOW to build it. ce-work executes the plan. A prior brainstorm is useful context but never required — ce-plan works from any input: a requirements doc, a bug report, a feature idea, or a rough description.
When directly invoked, always plan. Never classify a direct invocation as "not a planning task" and abandon the workflow. If the input is unclear, ask clarifying questions or use the planning bootstrap (Phase 0.4) to establish enough context — but always stay in the planning workflow.
This workflow produces a durable implementation plan. It does not implement code, run tests, or learn from execution-time results. If the answer depends on changing code and seeing what happens, that belongs in ce-work, not here.
Interaction Method
When asking the user a question, use the platform's blocking question tool: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded), request_user_input in Codex, ask_user in Gemini, ask_user in Pi (requires the pi-ask-user extension). Fall back to numbered options in chat only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question.
Ask one question at a time. Prefer a concise single-select choice when natural options exist.
Feature Description
<feature_description> #$ARGUMENTS </feature_description>
If the feature description above is empty, ask the user: "What would you like to plan? Describe the task, goal, or project you have in mind." Then wait for their response before continuing.
If the input is present but unclear or underspecified, do not abandon — ask one or two clarifying questions, or proceed to Phase 0.4's planning bootstrap to establish enough context. The goal is always to help the user plan, never to exit the workflow.
IMPORTANT: All file references in the plan document must use repo-relative paths (e.g., `src/models/user.rb`), never absolute paths (e.g., `/Users/name/Code/project/src/models/user.rb`). This applies everywhere — implementation unit file lists, pattern references, origin document links, and prose mentions. Absolute paths break portability across machines, worktrees, and teammates.
Core Principles
1. Use requirements as the source of truth - If ce-brainstorm produced a requirements document, planning should build from it rather than re-inventing behavior. 2. Decisions, not code - Capture approach, boundaries, files, dependencies, risks, and test scenarios. Do not pre-write implementation code or shell command choreography. Pseudo-code sketches or DSL grammars that communicate high-level technical design are welcome when they help a reviewer validate direction — but they must be explicitly framed as directional guidance, not implementation specification. 3. Research before structuring - Explore the codebase, institutional learnings, and external guidance when warranted before finalizing the plan. 4. Right-size the artifact - Small work gets a compact plan. Large work gets more structure. The philosophy stays the same at every depth. 5. Separate planning from execution discovery - Resolve planning-time questions here. Explicitly defer execution-time unknowns to implementation. 6. Keep the plan portable - The plan should work as a living document, review artifact, or issue body without embedding tool-specific executor instructions. 7. Carry execution posture lightly when it matters - If the request, origin document, or repo context clearly implies test-first, characterization-first, or another non-default execution posture, reflect that in the plan as a lightweight signal. Do not turn the plan into step-by-step execution choreography. 8. Honor user-named resources - When the user names a specific resource — a CLI, MCP server, URL, file, doc link, or prior artifact — treat it as authoritative input, not a suggestion. Discover it if unknown (command -v, fetch, read) before assuming it's unavailable. Use it in place of generic alternatives. If it fails or doesn't exist, say so explicitly rather than silently substituting.
Plan Quality Bar
Every plan should contain:
- A clear problem frame and scope boundary
- Concrete requirements traceability back to the request or origin document
- Repo-relative file paths for the work being proposed (never absolute paths — see Planning Rules)
- Explicit test file paths for feature-bearing implementation units
- Decisions with rationale, not just tasks
- Existing patterns or code references to follow
- Enumerated test scenarios for each feature-bearing unit, specific enough that an implementer knows exactly what to test without inventing coverage themselves
- Clear dependencies and sequencing
A plan is ready when an implementer can start confidently without needing the plan to write the code for them.
Workflow
Phase 0: Resume, Source, and Scope
0.0 Resolve Output Mode
Determine OUTPUT_FORMAT before any other phase fires. Output mode is exclusive — the plan is written as either markdown (.md) OR HTML (.html), never both. Precedence: CLI arg > config > default (md), with a hard pipeline-mode override.
Read config (pre-resolved at skill load): !cat "$(git rev-parse --show-toplevel 2>/dev/null)/.compound-engineering/config.local.yaml" 2>/dev/null || echo '__NO_CONFIG__'
Resolution steps:
1. CLI arg. Scan $ARGUMENTS for a token starting with the literal prefix output:. If found, strip it from arguments before treating the remainder as the feature description, and match its value case-insensitively against md and html.
output:alone (no value) → no-op, fall through to step 2.output:<unknown>(e.g.,output:pdf) → drop the token, fall through to step 2, and remember to emit a one-line note above the post-generation menu after final resolution:Ignored unknown output: value '<value>' — using <resolved_format> instead.where<resolved_format>is the valueOUTPUT_FORMATactually resolved to after steps 2-4. Do not hardcodemdin the note — that misleads users when config has set HTML.
2. Config. If step 1 did not resolve and the pre-resolved YAML above has an active (non-commented) plan_output: key whose value matches md or html (case-insensitive), use it. Missing, invalid, or commented values fall through silently. Critical: lines starting with # are YAML comments and must be ignored — the shipped config template includes commented examples like # plan_output: html to document the option, and matching those as active settings would silently force HTML mode on every run without the user having opted in. 3. Default. Otherwise OUTPUT_FORMAT=md. 4. Pipeline override. When invoked from LFG or any disable-model-invocation context, force OUTPUT_FORMAT=md regardless of steps 1-3. ce-work and other automated downstream consumers parse markdown reliably; HTML in pipeline runs is unnecessary friction.
Token-parsing convention: only literal-prefix flag tokens (output:, mode:, delegate: where applicable) are consumed and stripped. Other <word>:<word> tokens — including conventional commit prefixes like feat:, fix:, chore: that may appear inside a feature description — pass through verbatim.
Load the format-rendering reference based on the resolved value. Section content is the same in either format; presentation differs. Both references are paired with references/plan-sections.md, which describes what the plan contains regardless of format.
- When
OUTPUT_FORMAT=md, readreferences/markdown-rendering.mdfor format principles. - When
OUTPUT_FORMAT=html, readreferences/html-rendering.mdfor format principles.
0.1 Resume Existing Plan Work When Appropriate
If the user references an existing plan file or there is an obvious recent matching plan in docs/plans/:
- Read it
- Confirm whether to update it in place or create a new plan
- If updating, revise only the still-relevant sections. Plans do not carry per-unit progress state — progress is derived from git by
ce-work, so there is no progress to preserve across edits
Deepen intent: The word "deepen" (or "deepening") in reference to a plan is the primary trigger for the deepening fast path. When the user says "deepen the plan", "deepen my plan", "run a deepening pass", or similar, the target document is a plan in docs/plans/, not a requirements document. Use any path, keyword, or context the user provides to identify the right plan. If a path is provided, verify it is actually a plan document. If the match is not obvious, confirm with the user before proceeding.
Words like "strengthen", "confidence", "gaps", and "rigor" are NOT sufficient on their own to trigger deepening. These words appear in normal editing requests ("strengthen that section about the diagram", "there are gaps in the test scenarios") and should not cause a holistic deepening pass. Only treat them as deepening intent when the request clearly targets the plan as a whole and does not name a specific section or content area to change — and even then, prefer to confirm with the user before entering the deepening flow.
Once the plan is identified and appears complete (all major sections present, implementation units defined):
- Routing is keyed on file extension first, then frontmatter. HTML plans (
.html) are always software plans — the html-rendering invariant forbids YAML frontmatter, so frontmatter absence is not a non-software signal for HTML. Treat the visible-header metadata (title, date) as the frontmatter equivalent. - `.html` plan: short-circuit to Phase 5.3 (Confidence Check and Deepening) in interactive mode. Never route to
references/universal-planning.mdbased on missing YAML. - `.md` plan WITH YAML frontmatter: short-circuit to Phase 5.3 in interactive mode.
- `.md` plan WITHOUT YAML frontmatter (non-software plans use a simple
# Titleheading withCreated:date instead): route toreferences/universal-planning.mdfor editing or deepening instead of Phase 5.3. Non-software plans do not use the software confidence check.
The Phase 5.3 short-circuit avoids re-running the full planning workflow and gives the user control over which findings are integrated.
Normal editing requests (e.g., "update the test scenarios", "add a new implementation unit", "strengthen the risk section") should NOT trigger the fast path — they follow the standard resume flow.
If the plan already has a deepened: YYYY-MM-DD frontmatter field and there is no explicit user request to re-deepen, the fast path still applies the same confidence-gap evaluation — it does not force deepening.
Resume preserves the existing artifact's format, except pipeline mode. When resuming an existing plan, the resume run writes back in whatever format the existing artifact uses — markdown if the existing file is .md, HTML if it is .html — so a resume doesn't silently change the artifact shape. Explicit output: arguments on this run override (e.g., resuming an .html plan with output:md switches the artifact to markdown). Pipeline mode (LFG, any disable-model-invocation context) always wins per Phase 0.0: even when resuming an existing .html plan, pipeline runs force OUTPUT_FORMAT=md so downstream automation receives the markdown shape it expects. The resume rewrites the markdown file at the parallel path (<plan-basename>.md) and the original .html is left in place untouched.
0.1a Recognize Approach-Altitude Requests
Some requests are better answered one level up: produce a grounded approach-plan — a plan for how the deliverable will be made — and hold there, rather than zero-shotting the deliverable. This runs after Phase 0.1's resume and deepen fast paths (so "deepen the plan" and resume short-circuit first) and before Phase 0.1b's domain split (so the capability is domain-general — it applies to software and knowledge-work alike).
Two entries, with very different gating:
Explicit (always honored, ungated). When the user asks for the approach itself — "plan for a plan", "plan the approach", "plan how you'll do X", "don't do it yet -- just plan how you'd approach it" — enter approach altitude and hold at the approach. Do NOT begin the deliverable. Key on language that asks for the approach to producing something, not the something. This is a distinct signal from "deepen"/"strengthen" (the Phase 0.1 deepening fast path) and from a normal plan request.
Proactive (rare, conservative). When the user gives a plain request with no approach-language, offer an approach-plan only when both of these are clearly high:
- Method uncertainty — the core approach is genuinely unsettled: competing methodologies that would yield different deliverables, unclear how disparate sources or constraints combine, or an outcome stated only at the value level ("something I can actually use"). This is not satisfied by a task whose core method is obvious but whose rollout, sequencing, scope, or ordering has routine variants (big-bang vs. incremental, batch order, phased vs. one-shot) — those are ordinary plan decisions the Phase 0.7 scoping synthesis already surfaces as call-outs, not method-uncertainty. A large or mechanical change (a 40-endpoint migration, a wide rename, a framework bump) is typically costly but method-obvious; cost alone never fires the offer.
- Cost of getting it wrong — the deliverable is expensive or slow to produce and a wrong approach wastes real effort (heavy inputs to process, a long synthesis, a large or risky change).
If either is low, stay silent and plan/do normally. When borderline, stay silent. Assess this from request shape and input metadata only — do not read the inputs yet (recon happens after the offer is accepted). When the offer does fire, it is a single dismissible line naming the specific signal (e.g., "Three heavy sources are about to get synthesized and you might want them weighted differently -- want my approach first, or should I just go?") — never a blocking question, never a ceremony. Because the explicit path above is always available, a missed offer is cheap; the failure mode to avoid is the new-hammer nag — opening turns with "want me to plan the approach first?" when the method is obvious.
Stay disjoint from the other approach surfaces (R16). An investigative or analytical request with no approach-language and not-both-signals-high is NOT an approach-altitude request — it must pass through this gate untouched to Phase 0.1b, where answer-seeking's plan-of-attack handles it; the gate's earlier position must not intercept it. "Deepen the plan" and resume are already short-circuited by Phase 0.1. The Phase 0.7 / 5.1.5 scoping synthesis and the Phase 5.3 deepening pass operate on a deliverable already committed to; approach altitude operates before that commitment. Full distinctions: references/approach-altitude.md.
On entry (explicit, or an accepted offer), read references/approach-altitude.md and follow it. Otherwise continue to Phase 0.1b unchanged.
0.1b Classify Task Domain
If the task asks to build, modify, refactor, deploy, or architect software (code, schemas, infrastructure), continue to Phase 0.2.
Classify by task-type, not topic. A request that merely references code, a repo, an API, or a database is not automatically software work: building or modifying code is software; investigating or analyzing it is an answer-seeking question. "How often does X star repos — is it a big deal?" or "how does our approach compare to Y?" route to references/universal-planning.md (answer-seeking), not the implementation-plan path.
If the domain is genuinely ambiguous (e.g., "plan a migration" with no other context), ask the user before routing.
Otherwise, read references/universal-planning.md and follow that workflow instead. Skip all subsequent phases. Named tools or source links don't change this routing — they're inputs, handled per Core Principle 8.
0.2 Find Upstream Requirements Document
Before asking planning questions, search docs/brainstorms/ for files matching *-requirements.md or *-requirements.html (ce-brainstorm emits whichever extension matches its resolved output format; both are valid upstream requirements docs and either may be carried as the plan's origin:).
Relevance criteria: A requirements document is relevant if:
- The topic semantically matches the feature description
- It was created within the last 30 days (use judgment to override if the document is clearly still relevant or clearly stale)
- It appears to cover the same user problem or scope
If multiple source documents match, ask which one to use using the platform's blocking question tool when available (see Interaction Method). Otherwise, present numbered options in chat and wait for the user's reply before proceeding.
0.3 Use the Source Document as Primary Input
If a relevant requirements document exists: 1. Read it thoroughly 2. Announce that it will serve as the origin document for planning 3. Carry forward all of the following:
- Problem frame
- Actors (A-IDs), Key Flows (F-IDs), and Acceptance Examples (AE-IDs) when present — preserve these as constraints that implementation units must honor
- Requirements and success criteria
- Scope boundaries (including "Deferred for later" and "Outside this product's identity" subsections when present)
- Key decisions and rationale
- Dependencies or assumptions
- Outstanding questions, preserving whether they are blocking or deferred
4. Use the source document as the primary input to planning and research 5. Reference important carried-forward decisions in the plan with (see origin: <source-path>) 6. Do not silently omit source content — if the origin document discussed it, the plan must address it even if briefly. Before finalizing, scan each section of the origin document to verify nothing was dropped.
If no relevant requirements document exists, planning may proceed from the user's request directly.
0.4 Planning Bootstrap (No Requirements Doc or Unclear Input)
If no relevant requirements document exists, or the input needs more structure:
- Assess whether the request is already clear enough for direct technical planning — if so, continue to Phase 0.5
- If the ambiguity is mainly product framing, user behavior, or scope definition, recommend
ce-brainstormas a suggestion — but always offer to continue planning here as well - If the user wants to continue here (or was already explicit about wanting a plan), run the planning bootstrap below
The planning bootstrap should establish:
- Problem frame
- Intended behavior
- Scope boundaries and obvious non-goals
- Success criteria
- Blocking questions or assumptions
Keep this bootstrap brief. It exists to preserve direct-entry convenience, not to replace a full brainstorm.
If the bootstrap uncovers major unresolved product questions:
- Recommend
ce-brainstormagain - If the user still wants to continue, require explicit assumptions before proceeding
If the bootstrap reveals that a different workflow would serve the user better:
- Bug-shaped prompt (user describes broken behavior — "fix the bug where X", error message, regression, "doesn't work"). Surface
ce-debugas a route-out option alongside continuing withce-planwhenever the bug surface is reachable (in cwd OR named repo found at another local path). Stay ince-plansilently when the named code can't be found anywhere local — paper-planning is the only useful output for unreachable surfaces.
When the bug is at another local path (not cwd):
- Announce the target explicitly before any cross-repo investigation: which path will be read AND where plan outputs will land (default: target repo's
docs/plans/, not cwd's). - Default: proceed from the target repo for both investigation and plan-write. The user can interrupt to redirect (switch context, paper-plan, abandon, etc.). No location menu — the announcement makes the cross-repo nature visible, and the user can speak up if they want something unusual.
- After announcing and proceeding, fire the standard ce-debug routing menu (continue with
ce-planvs switch toce-debug) — same shape as the in-cwd case. Cross-repo location and ce-debug skill routing are orthogonal decisions; do not merge them into a single question.
Reading code at another path is fine in principle — that's just file access. The harm to avoid is silent operation on the wrong repo, especially writing the plan doc somewhere it won't be discovered (a busyblock plan landing in cli-printing-press/docs/plans/ is a discoverability disaster). The announcement requirement makes the target visible; defaulting to the target repo for both investigation and outputs respects the user's stated intent (they named that repo); the orthogonal ce-debug menu keeps the skill-choice question clean.
The accessibility classification is conservative and may under-suggest in monorepos, dependency bugs, or after renames. Users can always invoke /ce-debug manually.
Headless mode: skip the ce-debug suggestion menu entirely; default to continuing with /ce-plan (the user's explicit invocation). There is no synchronous user to resolve a route-out choice, and auto-routing to ce-debug would change the skill mid-flight without authorization.
- Clear task ready to execute (known root cause, obvious fix, no architectural decisions) — suggest
ce-workas a faster alternative alongside continuing with planning. The user decides.
0.5 Classify Outstanding Questions Before Planning
If the origin document contains Resolve Before Planning or similar blocking questions:
- Review each one before proceeding
- Reclassify it into planning-owned work only if it is actually a technical, architectural, or research question
- Keep it as a blocker if it would change product behavior, scope, or success criteria
If true product blockers remain:
- Surface them clearly
- Ask the user, using the platform's blocking question tool when available (see Interaction Method), whether to:
1. Resume ce-brainstorm to resolve them 2. Convert them into explicit assumptions or decisions and continue
- Do not continue planning while true blockers remain unresolved
0.6 Assess Plan Depth
Classify the work into one of these plan depths:
- Lightweight - small, well-bounded, low ambiguity
- Standard - normal feature or bounded refactor with some technical decisions to document
- Deep - cross-cutting, strategic, high-risk, or highly ambiguous implementation work
If depth is unclear, ask one targeted question and then continue.
0.7 Solo-Mode Scoping Synthesis
Surface call-outs to the user — the specific forks in scope or approach where user input materially changes the plan — so scope can be corrected before Phase 1 research is spent. Sub-agent dispatch (repo-research-analyst, learnings-researcher, etc.) is the expensive next step this phase guards against wasted effort on.
Fires only in solo invocation — when Phase 0.2 found no upstream brainstorm doc AND Phase 0.4 stayed in ce-plan (did not route to ce-debug, ce-work, or universal-planning) AND Phase 0.5 cleared (no unresolved blockers) AND not on Phase 0.1 fast paths (resume normal, deepen-intent). Each guard is an explicit conditional. Skip Phase 0.7 entirely when any guard fails — brainstorm-sourced invocations defer to Phase 5.1.5 instead.
Read `references/synthesis-summary.md` before composing the scoping synthesis. It carries the affirmability test, keep-test criteria, detail test, summary shape budgets, granularity rules, anti-patterns, revision-vs-confirmation discipline, doc-shape routing, soft-cut behavior, self-redirect support, the worked PII compression example, and full headless-mode routing — all required for a well-shaped synthesis.
Required gate output — do not skip; silent proceeding is not allowed. Compose an internal three-bucket scope draft (Stated / Inferred / Out of scope — internal thinking that feeds plan-body routing at Phase 5.2, not the chat output below). Derive call-outs (specific forks where user input materially changes the plan), then emit one of the two literal templates below in chat before continuing to Phase 1.
Synthesis is pre-plan-write. The agent does NOT yet know how plan-write will sequence the work. Do not claim PR count ("one PR"), commit/branch shape, effort or time estimates, Implementation Unit boundaries, or exact file paths in the synthesis. The synthesis surfaces decisions knowable at THIS point — for the solo variant, that's the user's request plus the Phase 0.4 bootstrap dialogue plus the agent's own internal three-bucket draft. Phase 1 research has not happened yet and there is no upstream brainstorm; do not claim grounding from either. Plan-write produces the rest. This rule holds even when the agent has formed plan-write opinions earlier in the session — those stay internal until plan-write.
Summary shape: the summary is a scope claim — what the plan will target, what it will not — at affirm-or-redirect level. NOT an enumeration of Implementation Units. Form is prose, bullets, or mix; tier budgets are ceilings, not targets (Lightweight 1-3 lines; Standard up to 3-5 lines or 2-4 bullets; Deep up to 4-6 lines or 3-6 bullets). 1-2 lines per bullet, conversational not documentary. Less is correct when there isn't more to say. See reference for keep test, detail test, and source-vocabulary discipline.
Do NOT enumerate the touch surface. Sentences like "The touch surface is...", "This plan touches...", "The implementation reaches into..." are plan-pitch leaks. File paths, module names, directory introductions, and per-file change descriptions belong in the plan body (Implementation Units at Phase 5.2), not the synthesis. The synthesis names what the plan targets, not where the code lives.
Pre-emit scans. Before emitting the synthesis, scan the output:
- Bare ID references (
AE\d+,R\d+,F\d+,A\d+,U\d+) → replace with plain names. - File paths (
path/like.md,path/like.py, etc.) → cut unless the path IS the topic of an explicit fork in the call-outs.
Tier guard on auto-proceed: the auto-proceed path (announce without waiting for confirmation) fires only when plan depth is Lightweight AND zero call-outs survive. Standard and Deep plans always fire the confirmation gate, even with zero call-outs — substance earns the checkpoint, not interaction history.
Confirmation template (Standard/Deep regardless of call-out count, or any tier with one or more call-outs surviving):
````text Based on your request and our brief discussion, here's the scope I'm proposing to plan against:
[scope claim — what the plan will target, what it will not; affirm-or-redirect level; NOT an enumeration of Implementation Units]
Call outs: (omit this header when zero forks survived the keep test)
- [decision-level fork in 1-2 lines: name the choice and optional one-clause trade-off in parens. NO multi-sentence rationale, NO "my default is X" pitch]
Confirm and I'll proceed to research, drawing on this scope. (You can also redirect to /ce-brainstorm if this is bigger than you initially thought — I'll stop here and load it for you.) ````
Wait for user confirmation before continuing to Phase 1.
Auto-proceed template (Lightweight with zero call-outs only):
````text Planning: [1-3 line scope claim]
No open decisions to weigh in on — proceeding to research. Interrupt if I have the scope wrong. ````
Then continue to Phase 1 without a blocking question.
Headless mode: internal draft is composed but stage 2 (chat-time call-outs) is skipped — no synchronous user to confirm to. Continue to Phase 1 research as normal. At plan-write time (Phase 5.2), Inferred bets from the internal draft route to a ## Assumptions section in the plan instead of Key Technical Decisions. See references/synthesis-summary.md Headless mode for the full routing.
Phase 1: Gather Context
1.1 Local Research (Always Runs)
Prepare a concise planning context summary (a paragraph or two) to pass as input to the research agents:
- If an origin document exists, summarize the problem frame, requirements, and key decisions from that document
- Otherwise use the feature description directly
- If
STRATEGY.mdexists, read it and include the relevant pieces (target problem, approach, active tracks) in the summary so downstream research and planning decisions are anchored to product strategy - If
CONCEPTS.mdexists at repo root, read it — its definitions are the canonical names for domain entities, named processes, and status concepts. Plan with those terms rather than synonyms.
Run these agents in parallel:
- Task ce-repo-research-analyst(Scope: technology, architecture, patterns. {planning context summary})
- Task ce-learnings-researcher(planning context summary)
Collect:
- Technology stack and versions (used in section 1.2 to make sharper external research decisions)
- Architectural patterns and conventions to follow
- Implementation patterns, relevant files, modules, and tests
- AGENTS.md guidance that materially affects the plan, with CLAUDE.md used only as compatibility fallback when present
- Institutional learnings from
docs/solutions/ - Product strategy context when
STRATEGY.mdis present — flag any plan decisions that pull away from the active tracks or the stated approach
Slack context (opt-in) — never auto-dispatch. Route by condition:
- Tools available + user asked: Dispatch
ce-slack-researcherwith the planning context summary in parallel with other Phase 1.1 agents. If the origin document has a Slack context section, pass it verbatim so the researcher focuses on gaps. Include findings in consolidation. - Tools available + user didn't ask: Note in output: "Slack tools detected. Ask me to search Slack for organizational context at any point, or include it in your next prompt."
- No tools + user asked: Note in output: "Slack context was requested but no Slack tools are available. Install and authenticate the Slack plugin to enable organizational context search."
1.1b Detect Execution Posture Signals
Decide whether the plan should carry a lightweight execution posture signal.
Look for signals such as:
- The user explicitly asks for TDD, test-first, or characterization-first work
- The origin document calls for test-first implementation or exploratory hardening of legacy code
- Local research shows the target area is legacy, weakly tested, or historically fragile, suggesting characterization coverage before changing behavior
When the signal is clear, carry it forward silently in the relevant implementation units.
Ask the user only if the posture would materially change sequencing or risk and cannot be responsibly inferred.
1.2 Decide on External Research
Based on the origin document, user signals, and local findings, decide whether external research adds value and, if so, what kind. Resolve this in three stages: explicit-request priority, intent classification, then the implicit signals below.
Stage 1 — An explicit request takes precedence. If the user prompt or the origin requirements document explicitly asks for external input — a signal that the answer lives outside the repo, such as competitor/prior-art comparison, "what should we borrow", "from the web", "best practices", "official docs", "alternatives to", a market scan, or naming a specific external technology to consult — external research is required, regardless of how strong local patterns look. The list is illustrative; key on the signal, not the exact phrase — any wording that clearly points outside the repo qualifies. The skip conditions below do not apply to an explicit request. The only thing that overrides it is an explicit opt-out ("no web research", "skip external research"): honor that, skip, and note it. Improvement or quality verbs ("improve", "make better") carry no external signal on their own and never trigger research by themselves.
Stage 2 — Classify the research intent (whenever external research will run, from Stage 1 or the implicit signals below) so Phase 1.3 routes correctly. Use this mechanical test, not a fixed phrase list:
- Implementation-guidance — the approach or technology is already settled; the question is how to build it well (best practices, version-specific docs, API constraints, known pitfalls, deprecations).
- Landscape / option-discovery — the question is what options or prior art exist (competitor scans, build-vs-buy, library/provider selection, prior art, market signals, cross-domain analogies).
- Mixed — both: discover an unsettled external option set first, then research the shortlisted choice for implementation guidance.
Stage 3 — Implicit signals decide the call when no explicit request fired.
Read between the lines. Pay attention to signals from the conversation so far:
- User familiarity — Are they pointing to specific files or patterns? They likely know the codebase well.
- User intent — Do they want speed or thoroughness? Exploration or execution?
- Topic risk — Security, payments, external APIs warrant more caution regardless of user signals.
- Uncertainty level — Is the approach clear or still open-ended?
Leverage ce-repo-research-analyst's technology context:
The ce-repo-research-analyst output includes a structured Technology & Infrastructure summary. Use it to make sharper external research decisions:
- If specific frameworks and versions were detected (e.g., Rails 7.2, Next.js 14, Go 1.22), pass those exact identifiers to ce-framework-docs-researcher so it fetches version-specific documentation
- If the feature touches a technology layer the scan found well-established in the repo (e.g., existing Sidekiq jobs when planning a new background job), lean toward skipping external research -- local patterns are likely sufficient
- If the feature touches a technology layer the scan found absent or thin (e.g., no existing proto files when planning a new gRPC service), lean toward external research -- there are no local patterns to follow
- If the scan detected deployment infrastructure (Docker, K8s, serverless), note it in the planning context passed to downstream agents so they can account for deployment constraints
- If the scan detected a monorepo and scoped to a specific service, pass that service's tech context to downstream research agents -- not the aggregate of all services. If the scan surfaced the workspace map without scoping, use the feature description to identify the relevant service before proceeding with research
Always lean toward external research when:
- The topic is high-risk: security, payments, privacy, external APIs, migrations, compliance
- The codebase lacks relevant local patterns -- fewer than 3 direct examples of the pattern this plan needs
- Local patterns exist for an adjacent domain but not the exact one -- e.g., the codebase has HTTP clients but not webhook receivers, or has background jobs but not event-driven pub/sub. Adjacent patterns suggest the team is comfortable with the technology layer but may not know domain-specific pitfalls. When this signal is present, frame the external research query around the domain gap specifically, not the general technology
- The user is exploring unfamiliar territory
- The technology scan found the relevant layer absent or thin in the codebase
- The plan's recommendations depend on a genuinely external, unsettled option set — which library, provider, or approach to adopt, or what competitors and prior art do — even when local implementation patterns are strong (intent: landscape). Bound this implicit landscape trigger by three gates: (a) the option set genuinely lives outside the repo, (b) the decision materially shapes the plan (a KTD, dependency, or architecture choice — not an incidental detail), and (c) no settled local or team choice already exists. Improvement verbs alone never satisfy this.
Skip external research when (only when Stage 1 found no explicit request — an explicit request is never skipped):
- The codebase already shows a strong local pattern -- multiple direct examples (not adjacent-domain), recently touched, following current conventions
- The user already knows the intended shape
- Additional external context would add little practical value
- The technology scan found the relevant layer well-established with existing examples to follow
When an explicit request did fire but a settled local or team choice already exists, narrow the research rather than skipping it — research the current pitfalls, docs, and practices for the chosen library/pattern instead of re-surveying the whole option set.
Announce the decision and the intent briefly before continuing. Examples:
- "Your codebase has solid patterns for this. Proceeding without external research."
- "This involves payment processing, so I'll research current best practices first (implementation-guidance)."
- "You asked what to borrow from competitors, so I'll run a landscape scan first (landscape/option-discovery)."
1.3 External Research (Conditional)
If Step 1.2 indicates external research is useful, dispatch by the intent classified in Stage 2, using the platform's subagent primitive (Agent/Task in Claude Code, spawn_agent in Codex, subagent in Pi). For ce-web-researcher, pass a focus hint plus the planning context summary and do not pass codebase content — it operates externally.
- Implementation-guidance — run in parallel:
- Task ce-best-practices-researcher(planning context summary)
- Task ce-framework-docs-researcher(planning context summary, with exact frameworks/versions from Phase 1.1 where available)
- Landscape / option-discovery — Task ce-web-researcher(focus hint, planning context summary). When the request targets projects on a code host (e.g., "competitors on GitHub"), name the discovery dimensions in the focus hint: project names and URLs, release recency and activity, CLI/UX shape, install path, docs and examples, plugin/extension surfaces, recurring issue themes, and license — treating star counts as a weak signal only.
- Mixed — sequential, not parallel: run
ce-web-researcherfirst to map the landscape and produce a shortlist; then runce-framework-docs-researcherand/orce-best-practices-researcheragainst the shortlisted technologies only when their details materially shape the plan.
Tool-unavailable handling. ce-web-researcher self-checks for web tools and stops if they are missing. Never block on this: if it reports research unavailable, or any researcher fails, warn and proceed, and carry the gap into Phase 1.4 so the plan records it honestly — especially when the user explicitly requested external research, where a silent skip would leave the plan looking evidence-based when it is not.
1.4 Consolidate Research
Summarize:
- Relevant codebase patterns and file paths
- Relevant institutional learnings
- Organizational context from Slack conversations, if gathered (prior discussions, decisions, or domain knowledge relevant to the feature)
- External references, prior art, competitor/landscape findings, and best practices, if gathered
- Related issues, PRs, or prior art
- Any constraints that should materially shape the plan
Land external findings in decisions, not an appendix. Any external research that ran must surface where it changes a choice — Key Technical Decisions rationale, Alternatives, Risks, or Sources & Research — not as a detached list with no bearing on the plan. If a finding shaped nothing, it was not load-bearing; do not pad the plan with it.
Mark whether external research was load-bearing. Record a single internal flag: did external findings materially shape a KTD, Alternative, Scope boundary, or Risk? This flag answers only that question — it does not gate whether research runs (Phase 1.2 owns that decision). Phase 5.3.2 reads it to decide whether to enter a confidence-scoring pass.
Record requested-but-unavailable. If the user explicitly requested external research but it could not run (web tools unavailable, researcher failed), state that in the plan as an assumption or open question rather than presenting the plan as externally grounded.
1.4b Reclassify Depth When Research Reveals External Contract Surfaces
If the current classification is Lightweight and Phase 1 research found that the work touches any of these external contract surfaces, reclassify to Standard:
- Environment variables consumed by external systems, CI, or other repositories
- Exported public APIs, CLI flags, or command-line interface contracts
- CI/CD configuration files (
.github/workflows/,Dockerfile, deployment scripts) - Shared types or interfaces imported by downstream consumers
- Documentation referenced by external URLs or linked from other systems
This ensures flow analysis (Phase 1.5) runs and the confidence check (Phase 5.3) applies critical-section bonuses. Announce the reclassification briefly: "Reclassifying to Standard — this change touches [environment variables / exported APIs / CI config] with external consumers."
1.5 Flow and Edge-Case Analysis (Conditional)
For Standard or Deep plans, or when user flow completeness is still unclear, run:
- Task ce-spec-flow-analyzer(planning context summary, research findings)
Use the output to:
- Identify missing edge cases, state transitions, or handoff gaps
- Tighten requirements trace or verification strategy
- Add only the flow details that materially improve the plan
Phase 2: Resolve Planning Questions
Build a planning question list from:
- Deferred questions in the origin document
- Gaps discovered in repo or external research
- Technical decisions required to produce a useful plan
For each question, decide whether it should be:
- Resolved during planning - the answer is knowable from repo context, documentation, or user choice
- Deferred to implementation - the answer depends on code changes, runtime behavior, or execution-time discovery
Ask the user only when the answer materially affects architecture, scope, sequencing, or risk and cannot be responsibly inferred. Use the platform's blocking question tool when available (see Interaction Method).
Do not run tests, build the app, or probe runtime behavior in this phase. The goal is a strong plan, not partial execution.
Phase 3: Structure the Plan
3.1 Title and File Naming
- Draft a clear, searchable title using conventional format such as
feat: Add user authenticationorfix: Prevent checkout double-submit - Determine the plan type:
feat,fix, orrefactor - Build the filename following the repository convention:
docs/plans/YYYY-MM-DD-NNN-<type>-<descriptive-name>-plan.md - Create
docs/plans/if it does not exist - Check existing files for today's date to determine the next sequence number (zero-padded to 3 digits, starting at 001)
- Keep the descriptive name concise (3-5 words) and kebab-cased
- Examples:
2026-01-15-001-feat-user-authentication-flow-plan.md,2026-02-03-002-fix-checkout-race-condition-plan.md - Avoid: missing sequence numbers, vague names like "new-feature", invalid characters (colons, spaces)
3.2 Stakeholder and Impact Awareness
For Standard or Deep plans, briefly consider who is affected by this change — end users, developers, operations, other teams — and how that should shape the plan. For cross-cutting work, note affected parties in the System-Wide Impact section.
3.3 Break Work into Implementation Units
Break the work into logical implementation units. Each unit should represent one meaningful change that an implementer could typically land as an atomic commit.
Good units are:
- Focused on one component, behavior, or integration seam
- Usually touching a small cluster of related files
- Ordered by dependency
- Concrete enough for execution without pre-writing code
Avoid:
- 2-5 minute micro-steps
- Units that span multiple unrelated concerns
- Units that are so vague an implementer still has to invent the plan
Each unit carries a stable plan-local U-ID assigned in Phase 3.5 (U1, U2, …). U-IDs survive reordering, splitting, and deletion: new units take the next unused number, gaps are fine, and existing IDs are never renumbered. This lets ce-work reference units unambiguously across plan edits.
3.4 High-Level Technical Design
When the plan's technical approach has shape that prose alone doesn't carry well — architecture across components, sequencing across processes, state machines, branching gates, lifecycles, quantitative comparisons — include a High-Level Technical Design section that conveys the shape. The exact form (component diagram, sequence, swim lane, flowchart, state machine, decision matrix, pseudo-code grammar, bar chart for sizing concerns) is the agent's call per artifact — pick what makes the content land fastest for the reader.
See references/plan-sections.md for the section catalog including HTD's "include when material" criterion. See the format-rendering reference loaded at Phase 0.0 for how visualizations render in the target format (mermaid in markdown, inline SVG in HTML — with the layout-legibility principles around halo, contrast, and label placement when in HTML).
When the plan's approach is a one-paragraph pattern application that prose conveys directly, skip the section. The presence of HTD should earn its keep with content that genuinely benefits from visualization.
Plan diagrams render authoritative content alongside the prose — they are not "directional sketches." Do not add hedging captions like "directional guidance for review, not implementation specification" to plan diagrams; the prose-is-authoritative rule already governs disagreement, and the hedging weakens the diagram unnecessarily.
3.4b Output Structure (Optional)
For greenfield plans that create a new directory structure (new plugin, service, package, or module), include an ## Output Structure section with a file tree showing the expected layout. This gives reviewers the overall shape before diving into per-unit details.
When to include it:
- The plan creates 3+ new files in a new directory hierarchy
- The directory layout itself is a meaningful design decision
When to skip it:
- The plan only modifies existing files
- The plan creates 1-2 files in an existing directory — the per-unit file lists are sufficient
The tree is a scope declaration showing the expected output shape. It is not a constraint — the implementer may adjust the structure if implementation reveals a better layout. The per-unit **Files:** sections remain authoritative for what each unit creates or modifies.
3.5 Define Each Implementation Unit
Each unit is a level-3 heading carrying a stable U-ID prefix matching the format used for R/A/F/AE in requirements docs: ### U1. [Name]. Number sequentially within the plan starting at U1. Do not render units as bulleted list items or prefix them with - [ ] / - [x] checkbox markers. List-based unit titles fragment in every standard renderer because the per-unit fields (**Goal:**, **Files:**, **Approach:**, etc.) are written flush-left, which terminates CommonMark list continuation and detaches the fields from the unit they describe. Headings render correctly everywhere, are the right semantic match for sections containing multi-block content, and give each unit an anchor link. The plan is a decision artifact; execution progress is derived from git by ce-work rather than stored in the plan body.
Stability rule. Once assigned, a U-ID is never renumbered. Reordering units leaves their IDs in place (e.g., U1, U3, U5 in their new order is correct; renumbering to U1, U2, U3 is not). Splitting a unit keeps the original U-ID on the original concept and assigns the next unused number to the new unit. Deletion leaves a gap; gaps are fine. This rule matters most during deepening (Phase 5.3), which is the most likely accidental-renumber vector.
For each unit, include:
- Goal - what this unit accomplishes
- Requirements - which requirements or success criteria it advances (cite R-IDs, and A/F/AE IDs when origin supplies them)
- Dependencies - what must exist first (cite by U-ID, e.g., "U1, U3")
- Files - repo-relative file paths to create, modify, or test (never absolute paths)
- Approach - key decisions, data flow, component boundaries, or integration notes
- Execution note - optional, only when the unit benefits from a non-default execution posture such as test-first or characterization-first
- Technical design - optional pseudo-code or diagram when the unit's approach is non-obvious and prose alone would leave it ambiguous. Frame explicitly as directional guidance, not implementation specification
- Patterns to follow - existing code or conventions to mirror
- Test scenarios - enumerate the specific test cases the implementer should write, right-sized to the unit's complexity and risk. Consider each category below and include scenarios from every category that applies to this unit. A simple config change may need one scenario; a payment flow may need a dozen. The quality signal is specificity — each scenario should name the input, action, and expected outcome so the implementer doesn't have to invent coverage. For units with no behavioral change (pure config, scaffolding, styling), use
Test expectation: none -- [reason]instead of leaving the field blank. AE-link convention: when a test scenario directly enforces an origin Acceptance Example, prefix it withCovers AE<N>.(orCovers F<N> / AE<N>.). This is sparse-by-design — most test scenarios are finer-grained than AEs and do not link. Do not force AE links onto tests that only cover lower-level implementation details. - Happy path behaviors - core functionality with expected inputs and outputs
- Edge cases (when the unit has meaningful boundaries) - boundary values, empty inputs, nil/null states, concurrent access
- Error and failure paths (when the unit has failure modes) - invalid input, downstream service failures, timeout behavior, permission denials
- Integration scenarios (when the unit crosses layers) - behaviors that mocks alone will not prove, e.g., "creating X triggers callback Y which persists Z". Include these for any unit touching callbacks, middleware, or multi-layer interactions
- Verification - how an implementer should know the unit is complete, expressed as outcomes rather than shell command scripts
Every feature-bearing unit should include the test file path in **Files:**.
Use Execution note sparingly. Good uses include:
Execution note: Start with a failing integration test for the request/response contract.Execution note: Add characterization coverage before modifying this legacy parser.Execution note: Implement new domain behavior test-first.
Do not expand units into literal RED/GREEN/REFACTOR substeps.
3.6 Keep Planning-Time and Implementation-Time Unknowns Separate
If something is important but not knowable yet, record it explicitly under deferred implementation notes rather than pretending to resolve it in the plan.
Examples:
- Exact method or helper names
- Final SQL or query details after touching real code
- Runtime behavior that depends on seeing actual test failures
- Refactors that may become unnecessary once implementation starts
3.7 Anti-Expansion: Tangential Cleanup and Scope Creep Go to Deferred
Distinct from 3.6 (which is about unknowns at plan time): 3.7 is about known but tangential work that the agent notices while planning but that falls outside the user's confirmed scope. When research surfaces an adjacent refactor, a "while we're here" cleanup, or a scope-adjacent nice-to-have ("we could also add rate limiting"), route it to the existing ### Deferred to Follow-Up Work subsection in Scope Boundaries (Phase 4.2 Core Plan Template), not into active Implementation Units.
This reinforces the synthesis discipline established at Phase 0.7 / Phase 5.1.5 — the user's confirmed scope is what the active plan executes; everything else is deferred. Does NOT impose architectural bias on extend-vs-invent decisions within confirmed scope — that judgment stays with the agent (and is surfaced via the Phase 5.1.5 synthesis when material). The user's explicit ask overrides this default — if the user explicitly requested a refactor, it's in-scope, not deferred.
Phase 4: Write the Plan
NEVER CODE during this skill. Research, decide, and write the plan — do not start implementation.
Use one planning philosophy across all depths. Change the amount of detail, not the boundary between planning and execution.
4.1 Plan Depth Guidance
Lightweight
- Keep the plan compact
- Usually 2-4 implementation units
- Omit optional sections that add little value
Standard
- Use the full core template, omitting optional sections (including High-Level Technical Design) that add no value for this particular work
- Usually 3-6 implementation units
- Include risks, deferred questions, and system-wide impact when relevant
Deep
- Use the full core template plus optional analysis sections where warranted
- Usually 4-8 implementation units
- Group units into phases when that improves clarity
- Include alternatives considered, documentation impacts, and deeper risk treatment when warranted
4.1b Optional Deep Plan Extensions
For sufficiently large, risky, or cross-cutting work, add the sections that genuinely help:
- Alternative Approaches Considered
- Success Metrics
- Dependencies / Prerequisites
- Risk Analysis & Mitigation
- Phased Delivery
- Documentation Plan
- Operational / Rollout Notes
- Future Considerations only when they materially affect current design
Do not add these as boilerplate. Include them only when they improve execution quality or stakeholder alignment.
Alternatives Considered — what to vary. When this section is included, alternatives must differ on how the work is built: architecture, sequencing, boundaries, integration pattern, rollout strategy. Tiny implementation variants (which hash function, which serialization format) belong in Key Technical Decisions, not Alternatives. Product-shape alternatives (different actors, different core outcome, different positioning) belong in ce-brainstorm, not here — surface them back upstream rather than re-litigating product questions during planning.
4.2 Section Contract and Rendering
Compose the plan using two paired references:
references/plan-sections.md— the section contract. Describes what the plan contains: the outcome the plan must enable for downstream consumers, the hard floor (Summary, Problem Frame, Requirements, KTDs, Implementation Units), the include-when-material catalog (HTD, Scope Boundaries, Open Questions, System-Wide Impact, Risks & Dependencies, Acceptance Examples, Documentation/Operational Notes, Sources & Research), the agency-driven escape hatch (introduce new sections when content warrants), and the ID/content rules.- The format-rendering reference loaded at Phase 0.0 (
markdown-rendering.mdORhtml-rendering.md) — how to present the sections in the resolved output format.
The section catalog is the same regardless of format. Format-specific principles (table-vs-prose by content shape, ID prefix format, diagram rendering, etc.) live in the rendering reference.
Omit "include when material" sections that don't carry information for this specific plan. Filling a section with placeholder prose is worse than omitting it.
4.3 Planning Rules
- Horizontal rules (`---`) between top-level sections in Standard and Deep plans, mirroring the
ce-brainstormrequirements doc convention. Improves scannability of dense plans where many H2 sections sit close together. Omit for Lightweight plans where the whole doc fits on a single screen. - All file paths must be repo-relative — never use absolute paths like
/Users/name/Code/project/src/file.ts. Usesrc/file.tsinstead. Absolute paths make plans non-portable across machines, worktrees, and teammates. When a plan targets a different repo than the document's home, state the target repo once at the top of the plan (e.g.,**Target repo:** my-other-project) and use repo-relative paths throughout - Prefer path plus class/component/pattern references over brittle line numbers
- Do not include implementation code — no imports, exact method signatures, or framework-specific syntax
- Pseudo-code sketches and DSL grammars are allowed in the High-Level Technical Design section and per-unit technical design fields when they communicate design direction. Frame them explicitly as directional guidance, not implementation specification
- Mermaid diagrams are encouraged when they clarify relationships or flows that prose alone would make hard to follow — ERDs for data model changes, sequence diagrams for multi-service interactions, state diagrams for lifecycle transitions, flowcharts for complex branching logic
- Do not include git commands, commit messages, or exact test command recipes
- Do not expand implementation units into micro-step
RED/GREEN/REFACTORinstructions - Do not pretend an execution-time question is settled just to make the plan look complete
Phase 5: Final Review, Write File, and Handoff
5.1 Review Before Writing
Before finalizing, check:
- The plan does not invent product behavior that should have been defined in
ce-brainstorm - If there was no origin document, the bounded planning bootstrap established enough product clarity to plan responsibly
- Every major decision is grounded in the origin document or research
- Each implementation unit is concrete, dependency-ordered, and implementation-ready
- If test-first or characterization-first posture was explicit or strongly implied, the relevant units carry it forward with a lightweight
Execution note - Each feature-bearing unit has test scenarios from every applicable category (happy path, edge cases, error paths, integration) — right-sized to the unit's complexity, not padded or skimped
- Test scenarios name specific inputs, actions, and expected outcomes without becoming test code
- Feature-bearing units with blank or missing test scenarios are flagged as incomplete — feature-bearing units must have actual test scenarios, not just an annotation. The
Test expectation: none -- [reason]annotation is only valid for non-feature-bearing units (pure config, scaffolding, styling) - Deferred items are explicit and not hidden as fake certainty
- High-Level Technical Design presence audit (load-bearing). For each architecture trigger in Phase 3.4 that the plan content satisfies (3+ components with directed relationships, 3+ protocol steps, 3+ state machine states, lifecycle, 3+ decision points, 3+ data-flow stages, mode/flag combinations, DSL/API surface design, non-obvious single-component shape), verify a corresponding sketch/diagram is present in the High-Level Technical Design section. Count the firing triggers; count the sketches; the sketch count must be at least the count of distinct trigger categories that fired. Missing the section when a trigger fired, OR including the section but skipping a triggered sketch within it, is incomplete — return to Phase 3.4 and add the missing sketch. Token cost is not a valid reason to fail this check.
- If a High-Level Technical Design section is included, it uses the right medium for the work, carries the non-prescriptive framing, and does not contain implementation code (no imports, exact signatures, or framework-specific syntax)
- Per-unit technical design fields, if present, are concise and directional rather than copy-paste-ready
- If the plan creates a new directory structure, would an Output Structure tree help reviewers see the overall shape?
- If Scope Boundaries lists items that are planned work for a separate PR, issue, or repo, are they under
### Deferred to Follow-Up Workrather than mixed with true non-goals? - U-IDs are unique within the plan and follow the stability rule — no two units share an ID; reordering or splitting did not renumber existing units; gaps from deletions are preserved
- Would a visual aid (dependency graph, interaction diagram, comparison table) help a reader grasp the plan structure faster than scanning prose alone?
If the plan originated from a requirements document, re-read that document and verify:
- The chosen approach still matches the product intent
- Scope boundaries and success criteria are preserved
- Blocking questions were either resolved, explicitly assumed, or sent back to
ce-brainstorm - Every section of the origin document is addressed in the plan — scan each section to confirm nothing was silently dropped
- If origin supplies A/F/AE IDs: every origin R/F/AE that affects implementation is referenced in Requirements, a U-ID unit, test scenarios, verification, scope boundaries, or explicitly deferred. Actors are carried forward when they affect behavior, permissions, UX, orchestration, handoff, or verification. The standard is preservation of product intent, not mandatory ID spam — irrelevant origin IDs may be omitted
- If origin was Deep-product (origin contains an
Outside this product's identitysubsection): the plan's Scope Boundaries preserves the three-way split —Deferred for laterandOutside this product's identitycarried verbatim from origin,Deferred to Follow-Up Workreserved for plan-local implementation sequencing
5.1.5 Brainstorm-Sourced Scoping Synthesis
Surface plan-time call-outs to the user before Phase 5.2 commits the plan to disk — the latest cheap moment to catch plan-time scope errors. The brainstorm already validated WHAT to build; this phase surfaces HOW the plan will execute on the forks that matter.
Fires only when the plan was sourced from an upstream brainstorm doc (Phase 0.2 found a *-requirements.md or *-requirements.html match) AND not on Phase 0.1 fast paths (resume normal, deepen-intent). Skip Phase 5.1.5 in solo invocation — solo plans handled their synthesis in Phase 0.7.
Read `references/synthesis-summary.md` before composing the scoping synthesis. It carries the affirmability test, keep-test criteria, detail test, summary shape budgets, granularity rules, anti-patterns, revision-vs-confirmation discipline, doc-body reading rules, doc-shape routing, soft-cut behavior, self-redirect support, the worked PII compression example, and full headless-mode routing — all required for a well-shaped synthesis.
Required gate output — do not skip; silent proceeding is not allowed. Compose an internal three-bucket scope draft (Stated / Inferred / Out of scope — internal thinking that feeds plan-body routing at Phase 5.2, not the chat output below). Derive call-outs (specific forks where user input materially changes the plan), then emit one of the two literal templates below in chat before continuing to Phase 5.2.
Synthesis is pre-plan-write. The agent does NOT yet know how plan-write will sequence the work. Do not claim PR count ("one PR"), commit/branch shape, effort or time estimates, Implementation Unit boundaries, or exact file paths in the synthesis. The synthesis surfaces decisions knowable at THIS point (brainstorm + research + agent posture); plan-write produces the rest. This rule holds even when the agent has formed plan-write opinions earlier in the session — those stay internal until plan-write.
Summary shape: two paragraphs.
1. Brainstorm-scope restatement (1-2 sentences, prose). Restates the brainstorm's scope as orientation, in the brainstorm's own vocabulary. NOT an enumeration of Implementation Units, restated constraints, or listed acceptance examples — the user wrote those. 2. Plan-specific scoping decisions (prose, or bullets when multi-faceted). Scope-level commitments the agent made that the brainstorm did not: full brainstorm coverage vs. narrowed subset; adjacent refactors pulled in vs. held out; test scope at scenario level. Each item must be affirmable by the user without reading code. Form follows substance; tier budgets are ceilings, not targets (Lightweight 1-3 lines; Standard up to 3-5 lines or 2-4 bullets; Deep up to 4-6 lines or 3-6 bullets). 1-2 lines per bullet. Less is correct when there isn't more to say. See reference for keep test, detail test, and source-vocabulary discipline.
Do NOT enumerate the touch surface. Sentences like "The touch surface is...", "This plan touches...", "The implementation reaches into...", "Files modified include..." are plan-pitch leaks. File paths, module names, directory introductions, and per-file change descriptions belong in the plan body (Implementation Units at Phase 5.2), not the synthesis. The synthesis names what the plan targets, not where the code lives.
Pre-emit scans. Before emitting the synthesis, scan the output:
- Bare ID references (
AE\d+,R\d+,F\d+,A\d+,U\d+) → replace with plain names. - File paths (
path/like.md,path/like.py, etc.) → cut unless the path IS the topic of an explicit fork in the call-outs.
Tier guard on auto-proceed: the auto-proceed path (announce without waiting for confirmation) fires only when plan depth is Lightweight AND zero call-outs survive. Standard and Deep plans always fire the confirmation gate, even with zero call-outs — substance earns the checkpoint, not interaction history.
Confirmation template (Standard/Deep regardless of call-out count, or any tier with one or more call-outs surviving):
````text The brainstorm scopes [1-2 sentence restatement in the brainstorm's vocabulary as orientation; NOT an enumeration of Implementation Units, constraints, or acceptance examples].
This plan [plan-specific scoping decisions: full-brainstorm coverage vs. narrowed subset; adjacent refactors in or out; test scope at scenario level. NOT PR count, sequencing, IU lists, or file paths].
Call outs: (omit this header when zero forks survived the keep test)
- [plan-time fork in 1-2 lines: name the choice and optional one-clause trade-off in parens. NO multi-sentence rationale, NO "my default is X" pitch]
Confirm and I'll write the plan next, drawing on the brainstorm, research, and this synthesis. ````
Wait for user confirmation before continuing to Phase 5.2.
Auto-proceed template (Lightweight with zero call-outs only):
````text Planning [brief brainstorm-scope restatement] — [plan-specific shape in one clause].
No open decisions to weigh in on — proceeding to plan-write. Interrupt if I have the scope wrong. ````
Then continue to Phase 5.2 without a blocking question.
Headless mode: internal draft is composed but stage 2 (chat-time call-outs) is skipped — no synchronous user to confirm to. Proceed to Phase 5.2 plan-write. Inferred bets from the internal draft route to a ## Assumptions section in the plan instead of Key Technical Decisions. See references/synthesis-summary.md Headless mode for the full routing.
5.2 Write Plan File
REQUIRED: Write the plan file to disk before presenting any options.
Use the Write tool to save the complete plan to the resolved format's extension:
docs/plans/YYYY-MM-DD-NNN-<type>-<descriptive-name>-plan.<md|html>Extension follows OUTPUT_FORMAT from Phase 0.0 — .md when markdown, .html when HTML. Sequence number NNN is derived from existing plan files in docs/plans/ regardless of extension (count both .md and .html) to ensure unique daily ordering.
Compose the plan using the content from references/plan-sections.md and the format-specific principles from the rendering reference loaded at Phase 0.0 (markdown-rendering.md OR html-rendering.md).
Write tight. A section being material is not license to pad it. Hold every kept section to the prose-economy discipline in references/plan-sections.md: one idea per sentence, a requirement or unit is intent plus at most one qualifier, defer forks to Open Questions rather than specifying both arms, resolve superseded text in place rather than stacking strata. Before declaring the plan written, run the named test there — could the implementer find a contradiction in each section in one pass?
HTML composition timing. When OUTPUT_FORMAT=html, Phase 5.3 deepening runs before this write completes its final form, but ce-doc-review is skipped in HTML mode (its mutation mechanics are markdown-only today — see Phase 5.3.8 format gate in references/plan-handoff.md). The HTML artifact reflects deepening synthesis but not doc-review autofixes; this is a known gap until ce-doc-review gains HTML-aware mutation.
Confirm (use absolute path so the reference is clickable in modern terminals):
Plan written to <absolute path to plan>Pipeline mode: If invoked from an automated workflow such as LFG or any disable-model-invocation context, skip interactive questions. Make the needed choices automatically and proceed to writing the plan. Pipeline mode forces OUTPUT_FORMAT=md at Phase 0.0.
CONCEPTS.md gap-fill (only if the file already exists): If the plan body uses a domain term whose definition is missing from CONCEPTS.md, add the entry. Domain entities, named processes, and status concepts with project-specific meaning only — not file paths, class names, function signatures, or implementation decisions. CONCEPTS.md is a glossary, not a spec or catch-all. Follow the format set by existing entries. Apply silently. Skip entirely if CONCEPTS.md does not exist — creation is owned by ce-compound and ce-compound-refresh.
5.3 Confidence Check and Deepening
After writing the plan file, automatically evaluate whether the plan needs strengthening.
Two deepening modes:
- Auto mode (default during plan generation): Runs without asking the user for approval. The user sees what is being strengthened but does not need to make a decision. Sub-agent findings are synthesized directly into the plan.
- Interactive mode (activated by the re-deepen fast path in Phase 0.1): The user explicitly asked to deepen an existing plan. Sub-agent findings are presented individually for review before integration. The user can accept, reject, or discuss each agent's findings. Only accepted findings are synthesized into the plan.
Interactive mode exists because on-demand deepening is a different user posture — the user already has a plan they are invested in and wants to be surgical about what changes. This applies whether the plan was generated by this skill, written by hand, or produced by another tool.
ce-doc-review and this confidence check are different:
- Use the
ce-doc-reviewskill when the document needs clarity, simplification, completeness, or scope control - This confidence check strengthens rationale, sequencing, risk treatment, and system-wide thinking when the plan is structurally sound but still needs stronger grounding
Pipeline mode: This phase always runs in auto mode in pipeline/disable-model-invocation contexts. No user interaction needed.
5.3.1 Classify Plan Depth and Topic Risk
Determine the plan depth from the document:
- Lightweight - small, bounded, low ambiguity, usually 2-4 implementation units
- Standard - moderate complexity, some technical decisions, usually 3-6 units
- Deep - cross-cutting, high-risk, or strategically important work, usually 4-8 units or phased delivery
Build a risk profile. Treat these as high-risk signals:
- Authentication, authorization, or security-sensitive behavior
- Payments, billing, or financial flows
- Data migrations, backfills, or persistent data changes
- External APIs or third-party integrations
- Privacy, compliance, or user data handling
- Cross-interface parity or multi-surface behavior
- Significant rollout, monitoring, or operational concerns
5.3.2 Gate: Decide Whether to Deepen
- Lightweight plans usually do not need deepening unless they are high-risk
- Standard plans often benefit when one or more important sections still look thin
- Deep or high-risk plans often benefit from a targeted second pass
- Thin local grounding override: If Phase 1.2 triggered external research because local patterns were thin (fewer than 3 direct examples or adjacent-domain match), always proceed to scoring regardless of how grounded the plan appears. When the plan was built on unfamiliar territory, claims about system behavior are more likely to be assumptions than verified facts. The scoring pass is cheap — if the plan is genuinely solid, scoring finds nothing and exits quickly
- Load-bearing external research override: If Phase 1.4 marked external research as load-bearing (it materially shaped a KTD, Alternative, Scope boundary, or Risk), always proceed to scoring — even when local implementation patterns are strong. A landscape or prior-art finding can shape recommendations the local codebase cannot verify, and the thin-grounding override above would miss it. This enters the scoring pass only; it does not force deepening
If the plan already appears sufficiently grounded and neither the thin-grounding nor the load-bearing-external-research override applies, report "Confidence check passed — no sections need strengthening", then load `references/plan-handoff.md` now and execute 5.3.8 → 5.3.9 → 5.4 in sequence. Document review is mandatory for markdown plans — do not skip it because the confidence check passed. The two tools catch different classes of issues. For HTML plans (OUTPUT_FORMAT=html), the plan-handoff 5.3.8 format gate skips ce-doc-review since its mutation mechanics are markdown-only today; the menu summary surfaces that limitation explicitly.
5.3.3–5.3.7 Deepening Execution
When deepening is warranted, read references/deepening-workflow.md for confidence scoring checklists, section-to-agent dispatch mapping, execution mode selection, research execution, interactive finding review, and plan synthesis instructions. Execute steps 5.3.3 through 5.3.7 from that file, then return here for 5.3.8.
5.3.8–5.4 Document Review, Final Checks, and Post-Generation Options
STOP. Load `references/plan-handoff.md` now before continuing. It carries the full instructions for 5.3.8 (document review), 5.3.9 (final checks and cleanup), and 5.4 (post-generation handoff, including the Proof HITL flow, post-HITL re-review, and Issue Creation branching). This load is non-optional — without it, the agent renders the post-generation menu, captures the user's selection, and stops without firing the routed action. Document review at 5.3.8 runs unconditionally for OUTPUT_FORMAT=md regardless of whether the confidence check already ran; for OUTPUT_FORMAT=html, plan-handoff's 5.3.8 format gate skips ce-doc-review because its mutation mechanics are markdown-only today. The default mode for markdown is headless (mode:headless) — safe_auto fixes apply silently, remaining findings surface contextually above the menu, and a deeper interactive review is opt-in via free-form prompt.
After document review and final checks, print a one-line summary of the headless review state above the menu (e.g., Doc review applied 3 fixes. 2 decisions, 1 proposed fix, 4 FYI observations remain (1 at P1).; for HTML plans where 5.3.8 was skipped, print Doc review skipped — ce-doc-review is markdown-only today; the HTML plan was not reviewed.), then present the menu. The menu has 5 options when actionable findings remain (proposed_fixes_count + decisions_count > 0) and 4 options otherwise — including the FYI-only case AND the HTML-skip case (skipped_reason: output_format_html), both of which hide option 2 because ce-doc-review's walkthrough is gated to actionable markdown findings and would have nothing valid to walk through. See references/plan-handoff.md for the full rule. Render the 5-option menu as a numbered list in chat per the AGENTS.md narrow exception for legitimate option overflow, with the hint "Pick a number or describe what you want." On platforms whose blocking question tool has no option cap (Codex request_user_input, Pi ask_user), use the platform's blocking tool; when that tool is unavailable or errors (e.g., Codex edit modes where request_user_input is not exposed), fall back to the same numbered-list-in-chat rendering with the "Pick a number or describe what you want." hint. The 4-option case routes through the platform's blocking tool normally (AskUserQuestion in Claude Code — call ToolSearch with select:AskUserQuestion first if its schema isn't loaded), with the same numbered-list-in-chat fallback when no blocking tool is available or the call errors. Never silently skip the question.
Question: "Plan ready at <absolute path to plan>. What would you like to do next?" (use absolute path so the reference is clickable in modern terminals)
Options. Option 4's label matches the artifact's format. Under exclusive output mode, exactly one of "Open in Proof" or "Open in browser" applies per run — OUTPUT_FORMAT=md shows Proof; OUTPUT_FORMAT=html shows browser. Proof operates on markdown and cannot ingest HTML; the browser option opens the local .html file. Render the option matching the format produced this run.
1. Start `/ce-work` (recommended) - Begin implementing this plan in the current session 2. Run deeper doc review - Walk through the remaining findings interactively (full ce-doc-review walkthrough) 3. Create Issue - Create a tracked issue from this plan in your configured issue tracker (GitHub or Linear) 4. Open in Proof (web app) — review and comment to iterate with the agent - Open the doc in Every's Proof editor, iterate with the agent via comments, or copy a link to share with others. Render only when `OUTPUT_FORMAT=md`. 4. Open in browser - Open the HTML plan file locally for review and sharing. Render only when `OUTPUT_FORMAT=html`. 5. Done for now - Pause; the plan file is saved and can be resumed later
Routing. Act on the user's selection — do not just announce it. Elaborate sub-flows (Proof HITL state machine, Issue Creation tracker detection, post-HITL resync) live in references/plan-handoff.md.
- Start `/ce-work` — Invoke the
ce-workskill via the platform's skill-invocation primitive (Skillin Claude Code,Skillin Codex, the equivalent on Gemini/Pi), passing the plan path as the skill argument. Do not merely tell the user to type/ce-work— fire the invocation now so the plan executes in this session. - Run deeper doc review — Re-invoke the
ce-doc-reviewskill on the plan path withoutmode:headlessso the interactive routing question and walkthrough fire. After it returns, re-render this menu with refreshed counts so the user can pick a next-stage action. - Create Issue — Detect the project tracker (
ghfor GitHub,linearfor Linear) and create the issue from the plan file as described under "Issue Creation" inreferences/plan-handoff.md. After creation, display the issue URL and ask whether to proceed to/ce-workvia the platform's blocking question tool. - Open in Proof (web app) — review and comment to iterate with the agent — Load the
ce-proofskill in HITL-review mode with the plan file assource file, the plan title asdoc title, identityai:compound-engineering/Compound Engineering, and recommended next step/ce-work. Then follow the post-HITL resync logic inreferences/plan-handoff.md, which handles the fource-proofreturn statuses, re-runsce-doc-reviewafter material edits, and falls back gracefully on upload failure. - Open in browser — Display the absolute path to the
.htmlplan file so the user can open it locally. Where the platform exposes a browser-opening primitive (e.g.,openon macOS,xdg-openon Linux,starton Windows), the agent may use it; otherwise print the absolute path and let the user open it. Do not invokece-workfrom this option — the user picked HTML for review/sharing, not handoff. - Done for now — Display a brief confirmation that the plan file is saved and end the turn. Do not start follow-up work without an explicit further user prompt.
If the user types free-form prompts targeting the findings (e.g., "review", "walk through", "deep review"), route as if they picked Run deeper doc review — fire the skill rather than looping back to the menu. For other free-text revisions, accept the input and loop back to this menu after applying the revision.
Completion check: This skill is not complete until the post-generation menu above has been presented, the user has selected an action, and the inline routing for that selection has been executed. Presenting the menu and stopping at the user's selection is not completion — fire the routed action.
Pipeline mode exception: In LFG or any disable-model-invocation context, skip the interactive menu and return control to the caller after the plan file is written, confidence check has run, and ce-doc-review has run in headless mode (per references/plan-handoff.md). Pipeline mode forces OUTPUT_FORMAT=md at Phase 0.0, so the 5.3.8 format gate never selects the HTML skip path in pipeline runs.
Approach Altitude
Loaded from SKILL.md Phase 0.1a when a request is answered one level up — produce a grounded approach-plan (a plan for how the deliverable will be made), hold at a checkpoint, then execute now or save for later. Entered explicitly ("plan for a plan") or via an accepted proactive offer. Domain-general: the deliverable may be a document, a synthesis, a study artifact, or a software implementation plan. The boundary this preserves is code vs. knowledge-work, not plan vs. execute — ce-plan never writes or runs code (Phase 4 / SKILL.md line 15); code execution always belongs to ce-work.
Stage 1: Light recon (cheap grounding)
The whole point of the approach-plan is to be specific enough to judge. Generic methodology ("read the book, extract themes, synthesize") is not worth approving. So before composing it, skim the provided inputs enough to ground the approach in specifics — not the full read; that is the deliverable's work, deferred to execution.
- Bound the recon per input type so the checkpoint stays cheap. Directional guidance, not a rule: for a PDF, section headers + first/last pages + a few sampled sections; for a long transcript, sampled spans plus topic shifts; for a codebase, entry points and the relevant module shape. Skim to locate what matters and how the pieces relate, then stop.
- Ground in specifics: name the concrete bridges the approach will make ("the transcript spends ~40 minutes on pricing, which maps to the book's chapter-3 framework — I'll connect them there"), not a generic recipe.
- Degrade gracefully. If the inputs are absent or arrive later, fall back to proposing from the request alone and flag the approach-plan as provisional/ungrounded — never block waiting for inputs, never emit generic methodology dressed as a plan.
- No process exhaust. The approach-plan reads as value to the user, not as an audit log of recon steps ("I skimmed the PDF, then sampled the transcript, then…"). Surface what you concluded, not the plumbing. (See the Veil of value in
references/universal-planning.md.)
Stage 2: Compose the approach-plan (chat-first)
Deliver the approach-plan in chat. It is file-optional — the user decides whether to persist it. Keep it scannable. Cover, right-sized to the request:
- How each input will be handled — what you'll mine from each, grounded in the recon.
- How they combine — the synthesis strategy / sequencing; this is usually the risky part and the most valuable thing to confirm.
- The shape of the deliverable — structure/outline of what executing this will produce.
- The forks worth confirming — the few decisions where the user's steer materially changes the result (e.g., weighting one source over another, depth vs. breadth, audience).
- Open questions — anything genuinely unresolved that the user should answer before execution.
This is not a software plan template (no implementation units / test scenarios) unless the deliverable itself is a software implementation plan — in which case "execute now / code" routes into the normal ce-plan flow (below) rather than composing the deliverable here.
Stage 3: Checkpoint
Hold at the approach. Use the platform's blocking question tool (AskUserQuestion in Claude Code — call ToolSearch with select:AskUserQuestion first if its schema isn't loaded; request_user_input in Codex; ask_user in Gemini/Pi). Fall back to numbered options in chat only when no blocking tool exists or the call errors — never silently skip.
Sequence orthogonal axes rather than cramming them into one menu (per the "split orthogonal decisions" rule and the 4-option cap):
1. First: "Execute now, or save for later?" 2. Then, only if executing now and the domain isn't already obvious: confirm code vs. knowledge-work deliverable. Offer to deepen the approach-plan as part of "save for later".
Stage 4: Route
Save for later. Persist the approach-plan to docs/plans/ so it survives. If the deliverable is non-code, write the marker (execution: knowledge-work, see references/plan-sections.md) at persist time — so a later ce-work invocation on the saved plan routes to the carve-out, not the code path. Offer to deepen it. Keep the plan agent-agnostic (no ce-work-specific choreography in the body) so any agent can execute it later.
Execute now -- code deliverable. The approach-plan's job is done; continue into the normal ce-plan flow (Phase 0.1b onward) to produce the implementation plan, then hand off to ce-work for the code. ce-plan never writes the code itself.
Execute now -- non-code deliverable. This is the knowledge-work path with no ce-work equivalent, so it routes to ce-work's carve-out:
1. Write the marker execution: knowledge-work into the plan frontmatter. 2. Persist the marked plan to docs/plans/ (the marker needs a file to live in so it can travel — R7's file-optional governs the user keeping a chat-only copy, but non-code execution forces a persist). 3. Fire the ce-work skill, passing the plan path, via the platform's skill-invocation primitive (Skill in Claude Code). Do not merely tell the user to run it — fire it so execution happens in this session.
ce-plan itself does not execute the deliverable in any path — it produces the approach-plan and hands off. The portable plan is also runnable by any other agent without ce-work.
Boundaries: not the other approach surfaces
Three in-chat "approach" mechanics already exist. Approach altitude is separate but coordinated — keep it disjoint by its distinguishing properties, not by vocabulary:
- Answer-seeking's plan-of-attack (
references/universal-planning.md): non-blocking (states the approach and proceeds immediately), discards its scaffold, produces a chat answer, and lives only in the non-software answer-seeking branch. Approach altitude is domain-general, holds at a checkpoint for a user decision, and produces a persistable, deepenable approach-plan. An investigative request with no approach-language is answer-seeking's, not this. - Scoping synthesis (Phase 0.7 / 5.1.5): a scope checkpoint for a deliverable already committed to — it confirms what the implementation plan will target. Approach altitude is an altitude checkpoint that decides whether to commit to the deliverable at all; it sits above the implementation plan, not inside producing one.
- Deepening (Phase 5.3): operates on a plan that already exists, strengthening it via confidence sub-agents. Approach altitude operates before any artifact exists. The "deepen" affordance offered at the approach-altitude checkpoint is the user optionally enriching the approach-plan — not the Phase 5.3 confidence pass.
Deepening Workflow
This file contains the confidence-check execution path (5.3.3-5.3.7). Load it only when the deepening gate at 5.3.2 determines that deepening is warranted.
5.3.3 Score Confidence Gaps
Use a checklist-first, risk-weighted scoring pass.
For each section, compute:
- Trigger count - number of checklist problems that apply
- Risk bonus - add 1 if the topic is high-risk and this section is materially relevant to that risk
- Critical-section bonus - add 1 for
Key Technical Decisions,Implementation Units,System-Wide Impact,Risks & Dependencies, orOpen QuestionsinStandardorDeepplans
Treat a section as a candidate if:
- it hits 2+ total points, or
- it hits 1+ point in a high-risk domain and the section is materially important
Choose only the top 2-5 sections by score. If deepening a lightweight plan (high-risk exception), cap at 1-2 sections.
If the plan already has a deepened: date:
- Prefer sections that have not yet been substantially strengthened, if their scores are comparable
- Revisit an already-deepened section only when it still scores clearly higher than alternatives
Section Checklists:
Requirements
- Requirements are vague or disconnected from implementation units
- Success criteria are missing or not reflected downstream
- Units do not clearly advance the traced requirements
- Origin requirements are not clearly carried forward
- Origin A/F/AE IDs (when supplied by the upstream brainstorm) are not preserved where planning decisions touch them, or are referenced inconsistently across Requirements, units, and test scenarios
Context & Research / Sources & References
- Relevant repo patterns are named but never used in decisions or implementation units
- Cited learnings or references do not materially shape the plan
- High-risk work lacks appropriate external or internal grounding
- Research is generic instead of tied to this repo or this plan
Key Technical Decisions
- A decision is stated without rationale
- Rationale does not explain tradeoffs or rejected alternatives
- The decision does not connect back to scope, requirements, or origin context
- An obvious design fork exists but the plan never addresses why one path won
Open Questions
- Product blockers are hidden as assumptions
- Planning-owned questions are incorrectly deferred to implementation
- Resolved questions have no clear basis in repo context, research, or origin decisions
- Deferred items are too vague to be useful later
High-Level Technical Design (when present)
- The sketch uses the wrong medium for the work
- The sketch contains implementation code rather than pseudo-code
- The non-prescriptive framing is missing or weak
- The sketch does not connect to the key technical decisions or implementation units
High-Level Technical Design (when absent) (Standard or Deep plans only)
- The work involves DSL design, API surface design, multi-component integration, complex data flow, or state-heavy lifecycle
- Key technical decisions would be easier to validate with a visual or pseudo-code representation
- The approach section of implementation units is thin and a higher-level technical design would provide context
Implementation Units
- Dependency order is unclear or likely wrong
- File paths or test file paths are missing where they should be explicit
- Units are too large, too vague, or broken into micro-steps
- Approach notes are thin or do not name the pattern to follow
- Test scenarios are vague (don't name inputs and expected outcomes), skip applicable categories (e.g., no error paths for a unit with failure modes, no integration scenarios for a unit crossing layers), or are disproportionate to the unit's complexity
- Feature-bearing units have blank or missing test scenarios (feature-bearing units require actual test scenarios; the
Test expectation: noneannotation is only valid for non-feature-bearing units) - Verification outcomes are vague or not expressed as observable results
- Existing U-IDs were renumbered after a unit was reordered, split, or deleted (U-IDs are stable: never renumber existing IDs; gaps from deletions are preserved; new units take the next unused number)
- A unit realizing an origin Key Flow does not cite the F-ID, or a unit enforcing an origin Acceptance Example does not cite the AE-ID, when origin supplies them
System-Wide Impact
- Affected interfaces, callbacks, middleware, entry points, or parity surfaces are missing
- Failure propagation is underexplored
- State lifecycle, caching, or data integrity risks are absent where relevant
- Integration coverage is weak for cross-layer work
Risks & Dependencies / Documentation / Operational Notes
- Risks are listed without mitigation
- Rollout, monitoring, migration, or support implications are missing when warranted
- External dependency assumptions are weak or unstated
- Security, privacy, performance, or data risks are absent where they obviously apply
Use the plan's own Context & Research and Sources & References as evidence. If those sections cite a pattern, learning, or risk that never affects decisions, implementation units, or verification, treat that as a confidence gap.
5.3.4 Report and Dispatch Targeted Research
Before dispatching agents, report what sections are being strengthened and why:
Strengthening [section names] — [brief reason for each, e.g., "decision rationale is thin", "cross-boundary effects aren't mapped"]For each selected section, choose the smallest useful agent set. Do not run every agent. Use at most 1-3 agents per section and usually no more than 8 agents total.
Use fully-qualified agent names inside Task calls.
Deterministic Section-to-Agent Mapping:
Requirements / Open Questions classification
ce-spec-flow-analyzerfor missing user flows, edge cases, and handoff gapsce-repo-research-analyst(Scope:architecture, patterns) for repo-grounded patterns, conventions, and implementation reality checks
Context & Research / Sources & References gaps
ce-learnings-researcherfor institutional knowledge and past solved problemsce-framework-docs-researcherfor official framework or library behaviorce-best-practices-researcherfor current external patterns and industry guidancece-web-researcherfor landscape/prior-art gaps — competitor patterns, market signals, or an unsettled external option set (which library/provider/approach) that recommendations depend on- Add
ce-git-history-analyzeronly when historical rationale or prior art is materially missing
Key Technical Decisions
ce-architecture-strategistfor design integrity, boundaries, and architectural tradeoffs- Add
ce-framework-docs-researcherorce-best-practices-researcherwhen the decision needs external grounding beyond repo evidence
High-Level Technical Design
ce-architecture-strategistfor validating that the technical design accurately represents the intended approach and identifying gapsce-repo-research-analyst(Scope:architecture, patterns) for grounding the technical design in existing repo patterns and conventions- Add
ce-best-practices-researcherwhen the technical design involves a DSL, API surface, or pattern that benefits from external validation
Implementation Units / Verification
ce-repo-research-analyst(Scope:patterns) for concrete file targets, patterns to follow, and repo-specific sequencing cluesce-pattern-recognition-specialistfor consistency, duplication risks, and alignment with existing patterns- Add
ce-spec-flow-analyzerwhen sequencing depends on user flow or handoff completeness
System-Wide Impact
ce-architecture-strategistfor cross-boundary effects, interface surfaces, and architectural knock-on impact- Add the specific specialist that matches the risk:
ce-performance-oraclefor scalability, latency, throughput, and resource-risk analysisce-security-sentinelfor auth, validation, exploit surfaces, and security boundary reviewce-data-integrity-guardianfor migrations, persistent state safety, consistency, and data lifecycle risks
Risks & Dependencies / Operational Notes
- Use the specialist that matches the actual risk:
ce-security-sentinelfor security, auth, privacy, and exploit riskce-data-integrity-guardianfor migrations, backfills, persistent data safety, constraints, transaction boundaries, and production data transformation risk (plan context — not the PR-reviewce-data-migration-reviewerpersona)ce-deployment-verification-agentfor rollout checklists, rollback planning, and launch verificationce-performance-oraclefor capacity, latency, and scaling concerns
Agent Prompt Shape:
For each selected section, pass:
- The scope prefix from the mapping above when the agent supports scoped invocation
- A short plan summary
- The exact section text
- Why the section was selected, including which checklist triggers fired
- The plan depth and risk profile
- A specific question to answer
Instruct the agent to return:
- findings that change planning quality
- stronger rationale, sequencing, verification, risk treatment, or references
- no implementation code
- no shell commands
5.3.5 Choose Research Execution Mode
Use the lightest mode that will work:
- Direct mode - Default. Use when the selected section set is small and the parent can safely read the agent outputs inline.
- Artifact-backed mode - Use only when the selected research scope is large enough that inline returns would create unnecessary context pressure.
Signals that justify artifact-backed mode:
- More than 5 agents are likely to return meaningful findings
- The selected section excerpts are long enough that repeating them in multiple agent outputs would be wasteful
- The topic is high-risk and likely to attract bulky source-backed analysis
If artifact-backed mode is not clearly warranted, stay in direct mode.
Artifact-backed mode uses a per-run OS-temp scratch directory. Create it once before dispatching sub-agents and capture its absolute path — pass that absolute path to each sub-agent so they write to it directly. Do not use .context/; the artifacts are per-run throwaway that are cleaned up when deepening ends (see 5.3.6b), matching the repo Scratch Space convention for one-shot artifacts. Do not pass unresolved shell-variable strings to sub-agents; they need the resolved absolute path.
SCRATCH_DIR="$(mktemp -d -t ce-plan-deepen-XXXXXX)"
echo "$SCRATCH_DIR"Refer to the echoed absolute path as <scratch-dir> throughout the rest of this workflow.
5.3.6 Run Targeted Research
Launch the selected agents in parallel using the execution mode chosen above. If the current platform does not support parallel dispatch, run them sequentially instead. Omit the mode parameter when dispatching so the user's configured permission settings apply.
Prefer local repo and institutional evidence first. Use external research only when the gap cannot be closed responsibly from repo context or already-cited sources.
If a selected section can be improved by reading the origin document more carefully, do that before dispatching external agents.
Direct mode: Have each selected agent return its findings directly to the parent. Keep the return payload focused: strongest findings only, the evidence or sources that matter, the concrete planning improvement implied by the finding.
Artifact-backed mode: For each selected agent, pass the absolute <scratch-dir> path captured earlier and instruct the agent to write one compact artifact file inside that directory, then return only a short completion summary. Each artifact should contain: target section, why selected, 3-7 findings, source-backed rationale, the specific plan change implied by each finding. No implementation code, no shell commands.
If an artifact is missing or clearly malformed, re-run that agent or fall back to direct-mode reasoning for that section.
If agent outputs conflict:
- Prefer repo-grounded and origin-grounded evidence over generic advice
- Prefer official framework documentation over secondary best-practice summaries when the conflict is about library behavior
- If a real tradeoff remains, record it explicitly in the plan
5.3.6b Interactive Finding Review (Interactive Mode Only)
Skip this step in auto mode — proceed directly to 5.3.7.
In interactive mode, present each agent's findings to the user before integration. For each agent that returned findings:
1. Summarize the agent and its target section — e.g., "The ce-architecture-strategist reviewed Key Technical Decisions and found:" 2. Present the findings concisely — bullet the key points, not the raw agent output. Include enough context for the user to evaluate: what the agent found, what evidence supports it, and what plan change it implies. 3. Ask the user using the platform's blocking question tool when available (see Interaction Method):
- Accept — integrate these findings into the plan
- Reject — discard these findings entirely
- Discuss — the user wants to talk through the findings before deciding
If the user chooses "Discuss", engage in brief dialogue about the findings and then re-ask with only accept/reject (no discuss option on the second ask). The user makes a deliberate choice either way.
When presenting findings from multiple agents targeting the same section, present them one agent at a time so the user can make independent decisions. Do not merge findings from different agents before showing them.
After all agents have been reviewed, carry only the accepted findings forward to 5.3.7.
If the user accepted no findings, report "No findings accepted — plan unchanged." Then proceed directly to Phase 5.4 (skip document-review and synthesis — the plan was not modified). This interactive-mode-only skip does not apply in auto mode; auto mode always proceeds through 5.3.7 and 5.3.8. No explicit scratch cleanup needed — $SCRATCH_DIR is OS temp and will be cleaned up by the OS; leaving it in place preserves the rejected agent artifacts for debugging.
If findings were accepted and the plan was modified, proceed through 5.3.7 and 5.3.8 as normal — document-review acts as a quality gate on the changes.
5.3.7 Synthesize and Update the Plan
Strengthen only the selected sections. Keep the plan coherent and preserve its overall structure.
In interactive mode: Only integrate findings the user accepted in 5.3.6b. If some findings from different agents touch the same section, reconcile them coherently but do not reintroduce rejected findings.
Deepening may tighten, not only grow. A section can be strengthened by cutting as well as adding — collapse multi-idea sentences, drop hedges, and delete superseded text outright rather than leaving it as strikethrough or stacking a separate "resolutions" layer on top of it. A shorter, contradiction-free section is a stronger one. This is distinct from "rewrite the entire plan from scratch" below, which stays forbidden.
Allowed changes:
- Tighten prose in a strengthened section: cut hedges, split sentences carrying more than one idea, and remove superseded text in place (version control holds the history)
- Clarify or strengthen decision rationale
- Tighten requirements trace or origin fidelity
- Reorder or split implementation units when sequencing is weak — but never renumber existing U-IDs. Reordering preserves U-IDs in their new order (e.g., U1, U3, U5 reordered is correct; renumbering to U1, U2, U3 is not). Splitting keeps the original U-ID on the original concept and assigns the next unused number to the new unit. Renumbering breaks ce-work blocker and verification references that were written against the original IDs
- Add missing pattern references, file/test paths, or verification outcomes
- Expand system-wide impact, risks, or rollout treatment where justified
- Reclassify open questions between
Resolved During PlanningandDeferred to Implementationwhen evidence supports the change - Strengthen, replace, or add a High-Level Technical Design section when the work warrants it and the current representation is weak
- Strengthen or add per-unit technical design fields where the unit's approach is non-obvious
- Add or update
deepened: YYYY-MM-DDin frontmatter when the plan was substantively improved
Do not:
- Add implementation code — no imports, exact method signatures, or framework-specific syntax. Pseudo-code sketches and DSL grammars are allowed
- Add git commands, commit choreography, or exact test command recipes
- Add generic
Research Insightssubsections everywhere - Rewrite the entire plan from scratch
- Invent new product requirements, scope changes, or success criteria without surfacing them explicitly
- Renumber existing U-IDs as part of reordering, splitting, deletion, or "tidying" the unit list. Deepening is the most likely accidental-renumber vector — preserve U-IDs even when the new order would look cleaner with sequential numbering
If research reveals a product-level ambiguity that should change behavior or scope:
- Do not silently decide it here
- Record it under
Open Questions - Recommend
ce-brainstormif the gap is truly product-defining
Markdown Rendering
This is a format-rendering reference — it describes how to render any artifact in markdown, independent of which skill is producing it.
It is paired with a section contract (plan-sections.md, brainstorm-sections.md, etc.) that describes what the artifact contains. This reference describes how markdown specifically presents it. The same content rendered by different skills shares the same markdown principles.
Hard invariants
These hold regardless of which skill produced the artifact.
- YAML frontmatter at the top of the file. Standard
---delimited block
containing the artifact's stable metadata (title, date, type, etc. — exact fields are per-skill, defined in the section contract).
- ASCII identifiers in anchors. Markdown headings auto-generate anchors
from the heading text. Keep headings ASCII so anchors are predictable (#implementation-units, not #implementación-units).
- Repo-relative paths for file references. Always. Never absolute paths
— they break portability across machines, worktrees, teammates.
- No HTML mixed in. Keep the markdown pure. No
<div>, no<details>,
no inline <style>. If a layout idea only works as HTML, defer it to the HTML rendering. Markdown stays markdown.
Format principles
These shape what "good" markdown looks like; the agent applies them per artifact based on content shape.
ID prefix format
Stable IDs (R, U, A, F, AE, KTD) appear as plain prefixes at the start of the bullet or heading — do NOT bold the prefix. The prefix is visually distinctive on its own; bolding it inflates visual noise.
- R1. The plan returns paginated sessions. ← right
- **R1.** The plan returns paginated sessions. ← wrong (bolded prefix)Same applies to unit headings: ### U1. Cloak detection in preflight contract.
Content shape: prose vs bullets vs tables
The same content can be rendered three ways; the agent picks per content shape, not by template default.
- Prose when the content has narrative flow (motivation, decision
rationale, problem framing). Bullets fragment narrative into disconnected pieces.
- Bullets when items share a parallel shape but each carries enough
prose to not fit a table cell.
- Tables when 5+ items share uniform structure (
ID + body,
name + value, decision + rationale, risk + mitigation). Tables scan faster at that scale and unlock additional columns (status, traceability, severity) that bullets can't accommodate cleanly.
The test: which shape would a reader scan fastest for this content? If items have parallel structure and 5+ instances, table. If items are 3-5 and each has a few lines of prose, bullets. If the content is a single narrative thought, prose.
Bold leader labels within bullets
When a bullet has substructure that benefits from named fields (Key Flows with Trigger / Actors / Steps / Outcome, Acceptance Examples with Covers / Given / When / Then), use bold leader labels at the start of nested bullets — not deeper heading levels.
- F1. Anonymous capture
- **Trigger:** Agent enters Step 2a with no session.
- **Actors:** A1, A2
- **Steps:** Preflight detects cloak; agent launches; capture proceeds.
- **Covered by:** R1, R2, R5This gives the bullet structure without needing H4/H5 headings that would clutter the doc and break TOC generation.
Section separators
For substantial artifacts, use horizontal rules (---) between top-level H2 sections. Omit for short docs where separators would dominate.
Tables for genuinely comparative info only
Use tables for the uniform-shape case in "Content shape" above. Don't use tables to render content lists that are really bullets — markdown tables are noisier in raw form and worse for diffs.
Section anatomy
How section types commonly render in markdown. These are patterns, not contracts — the agent picks the shape that fits the content.
- Summary / Problem Frame — prose paragraphs.
- Requirements — bullets with
R<N>.prefix. When requirements span
more than one concern, grouping under bold inline headers is the default shape, not optional polish (group by capability, not by discussion order); render a flat list only when every requirement is about the same thing. When requirements have status, traceability, or severity that warrant additional columns, escalate to a table.
- Implementation Units — H3 heading per unit with
U<N>.prefix.
Fields (Goal, Files, Patterns, Test Scenarios, Verification) render as bullets with bold leader labels, or as sub-headings if the field has multi-paragraph content.
- Key Technical Decisions — bullets with bold decision name + prose
rationale, or numbered KTD-N pattern when traceability matters.
- Key Flows / Acceptance Examples — bullets with bold leader labels
(Trigger / Actors / Steps / Outcome / Covers / Given-When-Then).
- Scope Boundaries — bullets, optionally split into "Deferred for
later" / "Outside this product's identity" sub-headings when the positioning distinction matters.
The agent picks more elaborate or simpler shapes based on what each specific artifact's content needs.
Diagrams
When the section contract calls for a diagram (architecture, sequence, flowchart, state machine, swim lane, data-flow), markdown renders it as a fenced mermaid block:
` ``mermaid
flowchart TB
A[Start] --> B{Decision}
B -->|yes| C[Action]
B -->|no| D[Other action]
` ``(TB direction default — keeps diagrams narrow in source view and in narrow rendered viewports.)
Markdown's diagram affordances are limited compared to HTML. For quantitative comparisons (bar charts, scatter plots) markdown has no native equivalent — use a table with the data and let prose or caption carry the interpretation. The richer visualization happens in the HTML rendering.
Inline code and code blocks
- Inline code for identifiers (variable names, function names,
flag names, file paths, IDs that aren't section anchors).
- Fenced code blocks with language tag for code, shell commands,
API request/response samples. Always specify the language for syntax highlighting and accessibility.
The flag `--cdp-url` accepts a URL.
` ``bash
browser-use --cdp-url http://localhost:9222
` ``No process exhaust
Engineering process metadata stays out of the artifact:
- No "captured at Phase X" notes
- No
## Next Stepspointing to the next skill - No italic provenance lines ("Brainstorm completed 2026-05-13")
- No engineering-flow shepherding ("Now read this file:", "Next, run that
command:")
This information belongs in commit messages, tool output, and agent transcripts — not in the artifact a reader returns to weeks later.
Frontmatter shape
Per-skill frontmatter fields are defined in each skill's section contract (plan-sections.md lists plan frontmatter; brainstorm-sections.md lists brainstorm frontmatter). Common rules:
- YAML at the top of the file, delimited by
---on its own line above
and below.
- Field names in lowercase snake_case (
created_at,topic, not
CreatedAt, Topic).
- No status / lifecycle field. Artifacts are point-in-time records
(decision or discovery), not tracked work items. Do not introduce a mutable status field or an active → completed lifecycle — whether the work shipped is derived from git, not stored in the doc.
- Stable across artifact revisions — never rename or repurpose a field.
Post-write audit
Before declaring the markdown file written, scan it for these common slips:
- All stable IDs are plain-prefix format, not bolded.
- No HTML elements mixed in.
- All file paths are repo-relative.
- Horizontal rule separators between H2s (for Standard / Deep artifacts).
- No process exhaust (Phase X notes, Next Steps pointers, provenance
lines).
- Tables only where 5+ uniform-shape items justify them.
- Frontmatter has all the per-skill required fields with reasonable values.
Plan Sections
This reference describes what makes a great implementation plan. It does NOT prescribe how the plan looks on the page — rendering is handled by the format-specific references (markdown-rendering.md, html-rendering.md).
The outcome
A great plan enables three audiences to act:
- The implementing agent (
ce-workor a human) starts from an informed
baseline — load-bearing decisions are named, research breadcrumbs orient their own investigation, unit boundaries are clear. The plan gives the implementer a starting point, not a substitute for their own investigation.
- The reviewer identifies the load-bearing decisions and the boundaries
of what's being changed in one pass.
- The future reader (anyone returning months later) traces why the work
was done, what shaped it, and where the artifacts live.
Sections earn their place by serving one of these audiences. Omit padding.
Decide whether a plan doc is warranted at all
Not every invocation of ce-plan should produce a plan document. For genuinely atomic work, the doc is ceremony — the implementer (whether ce-work or a human) can act directly without IDed units, KTDs, or Requirements as a checklist.
Bias toward producing a plan. The risk asymmetry favors writing one: a thin plan doc for small work is mild ceremony, but skipping a plan when one was warranted costs the implementer real time (reinvented decisions, lost unit boundaries, no IDed requirements to verify against). When unsure, write the plan.
Skip plan creation only when ALL of these hold:
- The work is atomic — fits in one commit, no meaningful unit boundaries
to break out independently.
- There are no design choices that constrain implementation — no
Key Technical Decisions worth recording. If the work needs the implementer to make a choice between two approaches, those approaches are KTDs and a plan is warranted.
- There are no scope boundaries worth pinning in writing — the work
scope is self-evident from the user's request.
- No upstream artifact (a brainstorm with R-IDs, an incident report,
a deferred-follow-up item from a prior plan) needs traceability through this plan.
Stress test the "looks atomic" case. Many requests look atomic at first glance but hide design decisions:
- "Add caching to this endpoint" — sounds atomic, but TTL, invalidation,
cache key shape, and backend selection are all KTDs. Write the plan.
- "Migrate from package A to package B" — sounds mechanical, but
semantic differences between the packages create migration KTDs. Write the plan.
- "Add rate limiting" — sounds small, but algorithm, scope, and
configurability are all KTDs. Write the plan.
vs. genuine skip cases:
- "Fix typo in README line 47" — atomic, no KTDs, skip the plan.
- "Rename `oldFn` to `newFn` across the repo" — mechanical, no design
choices, skip the plan.
- "Bump dependency X to v2.3.1" — mechanical, skip the plan (unless the
bump introduces breaking changes that warrant unit-by-unit migration).
When skipping the plan doc, the work proceeds directly to ce-work or to implementation, and any decisions made along the way land in the commit message or docs/solutions/ if they're worth carrying forward.
Hard floor
When a plan doc is warranted, these sections are present. They carry the contracts downstream consumers depend on.
- Summary — what the plan proposes, in 1-3 lines. Forward-looking; orients
the reader before they invest in detail.
- Problem Frame — why the work is being done. Backward-looking /
situational. May merge with Summary for compact plans where the motivation is a single sentence.
- Requirements (with stable R-IDs) — what must be true after the work
ships. Reviewer's checklist; downstream code review verifies against these.
- Key Technical Decisions (KTDs) — the load-bearing choices that constrain
implementation. Each entry is <decision>: <rationale>. Without these, the implementer can't tell which design choices are open and which are pinned.
- Implementation Units (with stable U-IDs) — the discrete units of work,
sized so each is independently landable. ce-work consumes these to execute. For trivial single-step plans the work may collapse into Summary prose and U-IDs may be omitted; this is rare.
Include when material
These sections are present when they carry information that isn't covered elsewhere. The test is not "is this a substantial plan?" — it is "does this specific plan have content this section would surface?" Filling a section with placeholder prose is worse than omitting it.
- High-Level Technical Design — include when the technical approach has
shape that prose alone doesn't carry well: architecture across components, sequencing across processes, state machines, branching gates. Visualizations (component topology, sequence, swim lane, flowchart, data-flow) typically live here. Skip when the approach is a one-paragraph pattern application that the prose itself conveys.
- Scope Boundaries — include when scope is contested, when there are
tempting non-goals worth naming explicitly, or when "deferred for later" needs distinguishing from "outside the product's identity." Skip when scope is obvious from Requirements alone.
- Open Questions — include when there are genuinely unresolved items that
block planning or implementation. Skip when the plan is complete; an empty "Open Questions: none" section signals false uncertainty.
- System-Wide Impact — include when the change affects cross-cutting
concerns (data lifecycles, auth boundaries, performance posture, cardinal rules, shared infrastructure). Skip for changes localized to one component where the impact is self-evident.
- Risks & Dependencies — include when there are real risks worth flagging
(external service changes, version pins under churn, behavioral assumptions worth highlighting) or material upstream dependencies. Skip for low-risk localized work.
- Acceptance Examples — include when any requirement has a state-dependent
or conditional shape ("When X, Y") where the prose alone leaves ambiguity about edge cases. Skip when all requirements are unconditional and unambiguous.
- Documentation / Operational Notes — include when documentation,
monitoring, runbooks, or rollout steps need explicit notes. Skip when the work is purely internal and uses existing operational scaffolding without modification.
- Sources / Research — surface the research that orients the implementer
or justifies load-bearing choices. The test: "if I were the implementer reading this cold, would this breadcrumb help me make better choices?" Yes → surface (code locations like services/convex/reports.ts:174-176, external docs, RFCs, constraints, prior plans — the category is inclusive, not enumerated). Process exhaust (reading the user's prompt, glancing at obvious entry points, restating prose) → omit. Surface inline next to the KTD or unit it justifies, or as a dedicated section — both shapes work.
Agent agency
The catalog is a floor, not a ceiling. When the plan's content doesn't fit any catalog section, introduce a new one — don't force the content into a section it doesn't belong in. Content drives section choices, not vice versa.
The agent also picks per artifact:
- Whether Problem Frame merges into Summary
- Sub-groupings (Requirements by capability, KTDs by component, Units phased
into milestones)
- How much detail each section carries
- Whether HTD has one diagram, several, or none — and whether visualizations
live in HTD or embedded in other sections
Prose economy
"Include when material" sizes which sections appear; this sizes how the kept prose reads. A section can be material and still be written loosely — the failure mode is a material section padded into a wall of text where contradictions hide and the implementing agent loses the thread. A deep plan earns length through coverage (more units, more traced requirements, real risks), never through wordiness around that coverage.
Hold every kept section to these:
- One idea per sentence. A Summary is a handful of sentences, not one
sentence with five semicolons and four parentheticals. A KTD's rationale is the load-bearing reason, not every reason.
- **A requirement or unit is one sentence of intent plus at most one
qualifier.** When it would specify two outcomes ("either A or B, the implementer decides"), state the intent and send the fork to Open Questions — don't write both arms in full inside the item.
- Cut hedges and intensifiers. "Critically", "deliberately", "explicitly",
"genuinely", "actually", "simply" carry nothing the implementer acts on.
- Prefer the verb to the nominalization. "Demote the grid", not "the
demotion of the grid is the deliberate change in this plan".
Precision is not padding: keep file paths, IDs, conditionals, and exact thresholds verbatim. Economy targets the connective tissue around them, never the precision itself.
Resolve in place; don't stratify. When deepening, a doc-review pass, or a later decision supersedes earlier text, rewrite or remove the original — don't leave it standing as strikethrough or stack a separate "resolutions" layer on top of it. Version control holds the history. Stacked strata double the reading surface and hide which text is live.
Named test, run before the plan is declared written: could the implementer find a contradiction in each section in one pass? A sentence carrying more than one parenthetical, or an item specifying two outcomes, fails the test — split it or defer it.
Plan metadata fields
Every plan carries a small set of stable metadata fields that downstream tooling depends on. The contract is format-independent: in markdown these fields appear as YAML frontmatter at the top of the file; in HTML they appear as visible header text (typically a <dl> of <dt>/<dd> pairs or a stats strip). Field names and semantics are the same across both formats so consumers can locate them without knowing which format produced the plan.
Required
- `title` — verbatim plan title. Matches the H1 (markdown) or document
<h1> (HTML) so file metadata and visible heading don't drift.
- `type` — conventional-commit-prefix-aligned classification (
feat,
fix, refactor, chore, docs, perf, test, etc.). Carries the intent the eventual commit message should reflect.
- `date` — creation date in ISO 8601 (
YYYY-MM-DD), ASCII digits only.
Plans carry no `status` field — a plan is a decision artifact, not a tracked work item. ce-work does not mutate the plan at ship time; whether a plan shipped is derived from git, not stored in the doc. Do not add a status field or an active → completed lifecycle.
Optional but well-known
These fields are not required, but when set they have fixed names and semantics so downstream tooling can rely on them:
- `origin` — repo-relative path to an upstream brainstorm requirements
doc (e.g., docs/brainstorms/2026-05-12-pagination-requirements.md). Set when planning from an upstream brainstorm; carried for traceability and re-resolved when ce-plan re-deepens. The HITL Proof flow uses origin to trace back to the source brainstorm.
- `deepened` — ISO 8601 date marking the first time the confidence
check substantively strengthened the plan. Presence affects Phase 0.1 resume fast-path logic (see references/deepening-workflow.md).
- `execution` — execution domain for downstream routing:
code
(the default when absent) or knowledge-work. ce-work's input triage reads this: a plan marked execution: knowledge-work routes to the non-code carve-out (read sources, synthesize, produce a deliverable — skipping the branch/test/commit/CI lifecycle); absent or code routes to the normal code path. Written by ce-plan's approach-altitude flow (references/approach-altitude.md) when a non-code deliverable is persisted for execution.
Field names are stable across plan revisions — never rename a field or repurpose its semantics. Agents composing new plans MUST use these exact names; adding new fields is fine, but renaming origin to source or date to created breaks the downstream consumers above.
ID and content rules
These apply regardless of rendering format.
- Stable IDs. R-IDs (Requirements), U-IDs (Implementation Units), A-IDs
(if Actors fire), F-IDs (if Flows fire), AE-IDs (if Acceptance Examples fire). IDs are stable across plan revisions — never renumber to "clean up gaps."
- Plain prefix.
R1.,U1.as bullet prefixes. Do not bold; the prefix
is visually distinctive on its own.
- Repo-relative paths. Always. Never absolute paths in plan content;
they break portability across machines, worktrees, teammates.
- No process exhaust. No "captured at Phase X" notes, no
## Next Steps
pointing to the next skill, no italic provenance lines. Engineering process metadata belongs in commit messages and tool output, not the artifact.
- Group Requirements by concern when they span distinct logical areas.
The trigger is distinct concerns, not item count — even four requirements benefit from grouping if they cover three different topics. Skip grouping only when all requirements are genuinely about the same thing; a long flat list is a smell that subgroups were missed. Group by capability (e.g., "Packaging", "Migration and compatibility", "Contributor workflow"), not by the order requirements were discussed. R-IDs stay continuous across groups (R1, R2 in the first group; R3, R4 in the second; never restart at R1 per group).
Rendering
The format-specific references describe how to render these sections in each output format:
- Markdown rendering:
references/markdown-rendering.md - HTML rendering:
references/html-rendering.md
This reference (plan-sections.md) is about WHAT the plan contains; rendering references are about HOW each format presents it. The plan is written in one format — markdown OR HTML, never both — based on the resolved output mode. The section catalog is the same regardless of format.