
Strategy Review
- 45 installs
- 74 repo stars
- Updated July 21, 2026
- existential-birds/beagle
Helps with ai & agent building tasks.
About
strategy-review is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted development.
- strategy-review
- AI & Agent Building
- AI-coding skill
Strategy Review by the numbers
- 45 all-time installs (skills.sh)
- Ranked #7,680 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
npx skills add https://github.com/existential-birds/beagle --skill strategy-reviewAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 45 |
|---|---|
| repo stars | ★ 74 |
| Last updated | July 21, 2026 |
| Repository | existential-birds/beagle ↗ |
What it does
Helps with ai & agent building tasks.
Files
Strategy Review
Pressure-test strategy documents to find where they'll break before reality does it for you. The primary job isn't to evaluate prose quality or check formatting — it's to find the gaps, hidden failure paths, and under-accounted risks that will kill the strategy in execution. A strategy that survives this review has a meaningfully better chance of surviving contact with the real world.
This skill complements the strategy-interview skill. Strategy-interview helps build a strategy through guided conversation; strategy-review subjects an existing strategy to rigorous adversarial evaluation using the same kernel framework (diagnosis, guiding policy, coherent actions) and bad-strategy filter.
What makes this different from generic feedback
Most strategy feedback falls into two useless categories: vague praise ("this is really thoughtful") or surface-level nitpicking ("consider adding a timeline"). Neither helps the author see whether their thinking actually holds up under stress.
This review does three things that generic feedback doesn't:
1. Tests structural integrity — whether the logical chain from diagnosis through guiding policy to coherent actions actually holds together, or whether the "therefore" between them is secretly an "and also." 2. Hunts for failure paths — not just what's wrong with the document, but what goes wrong in the real world because of what's missing. The slow drift scenario, the capability gap that stalls the load-bearing action, the competitor response nobody modeled, the political resistance the strategy pretends doesn't exist. 3. Surfaces invisible assumptions — every strategy is a bet, but most strategies don't name their bets. This review finds the load-bearing assumptions the author hasn't stated and maps the risk of each one being wrong.
Before you start
Read references/review-dimensions.md — it contains the seven evaluation dimensions and their criteria. This is the backbone of the review. Keep it in working memory throughout.
Hard gates (evidence-bound)
These steps are easy to rationalize without evidence; each gate has a pass condition before you advance.
1. Ratings (Step 3). Do not assign Strong, Adequate, or Weak for any dimension until you have at least one of: a quoted or section-referenced passage from the strategy inputs, a strategy-notes.md (or equivalent) cross-reference, or an explicit Missing note stating that no relevant passage exists. Pass: every dimension that is rated has one of those anchors on record (in chat, in dimension-ratings.md if using durable state, or inline in the final review). 2. Critical findings (Step 5). Do not list an item under Critical findings unless it ties to the same kind of anchor (quote, section ref, notes cross-ref, or explicit absence). Pass: no critical finding is only a generic critique with no tie to the document or notes. 3. Judge artifact mode. Order: finalize prose strategy-review.md → derive strategy-review.json from that prose → run the validation checklist in references/judge-artifact-schema.md → emit or save JSON only after the checklist passes (parses, required fields, score arithmetic). Pass: JSON validates and matches the prose labels and evidence. 4. Durable state (when `.beagle/.../reviews/.../` exists). Before you call the review complete, ensure review-state.md exists and its current_step and lens/kernel fields match what you actually did (Steps 1–5 and files on disk). Pass: ledger reflects the true step and inputs; if dimension-ratings.md or kernel-extraction.md were created, the final strategy-review.md does not contradict them.
Complementary review lenses
The kernel evaluation (diagnosis, guiding policy, coherent actions) is always the backbone. Four additional lenses load conditionally when the strategy document warrants them. Do not force them. Most reviews use one or two; some use none. A lens loads when the document's content triggers it — not as a mandatory checklist.
Lenses sharpen the existing seven dimensions rather than adding new ones. When a lens reveals a gap, the finding belongs under the relevant dimension. The lens is the tool that found the gap; the dimension is where it lives.
| Lens | When it loads | What it adds | Primary dimensions |
|---|---|---|---|
| McKinsey 7S (Execution Alignment) | Strategy requires organizational change — new structure, processes, cultural shift, or cross-functional coordination | Checks whether the strategy accounts for alignment across structure, systems, shared values, style, staff, and skills. The most common gap: the strategy changes direction but assumes the organization will follow. | Dim 3, Dim 6 |
| Balanced Scorecard (Objective Translation) | Success criteria are one-dimensional (usually financial only) or vague | Checks whether success is measurable across financial, customer, process, and learning perspectives. Financial metrics are lagging — the strategy needs leading indicators that signal problems before revenue confirms them. | Dim 7 |
| Porter's Five Forces (Competitive Pressure Audit) | Diagnosis addresses competitive dynamics — market position, pricing pressure, new entrants, substitution risk | Checks whether the diagnosis saw the full competitive picture, not just direct rivalry. Many strategies diagnose one force while ignoring the others. | Dim 1 |
| Hoshin Kanri (Strategy Deployment) | Strategy involves multiple organizational levels — actions need to cascade from executive intent to team execution | Checks whether actions survive translation from strategy to operations. Strategy fails at the translation layers — each level of the org loses fidelity. | Dim 3 |
Lens selection happens during Step 2 (reading and extracting the kernel). After identifying the kernel elements, mentally check each lens trigger. Load references/review-lenses.md for the relevant lenses and weave their pressure-test questions into the Step 3 dimension evaluation.
The reviewer uses lens thinking to find gaps — not lens vocabulary to lecture the user. A finding that says "the strategy changes direction but doesn't address whether the current org structure supports the new approach" is a 7S-informed finding. The user doesn't need to know it came from McKinsey's framework.
Review workflow
Step 1 — Gather the documents
Ask the user what they want reviewed. Typical inputs:
- `strategy-draft.md` + `strategy-notes.md` — the pair produced by strategy-interview. The notes file dramatically enriches the review because it contains assumptions, open questions, and the reasoning journey. Always ask for both if only one is provided.
- A standalone strategy document — any doc that claims to be a strategy. Could be a Google Doc paste, a PDF, a deck summary, or a markdown file.
- A slide deck or brief — extract the strategic claims and evaluate those. Don't critique slide design.
If the document is ambiguous about what it's trying to be (strategy? plan? vision? goals?), name it: "This reads more like a goals document than a strategy — it describes what you want to achieve but not the theory of how. Want me to review it as-is, or should we first identify the missing strategic elements?"
Check the working directory for existing strategy files before asking the user to provide them — they may have already been produced by a previous strategy-interview session.
Step 2 — Read and extract the kernel
Read the full document. Identify (or note the absence of) the three kernel elements:
1. Diagnosis — What does the document say is actually going on? Is there a clear statement of the challenge? 2. Guiding policy — What overall approach has been chosen? Does it make a directional choice? 3. Coherent actions — What concrete steps carry out the policy? Do they reinforce each other?
If the document uses different terminology (OKRs, pillars, strategic priorities, initiatives), map their concepts to the kernel. Don't force a vocabulary change — evaluate the thinking underneath the labels. A "strategic priority" that functions as a guiding policy should be evaluated as one.
Also note:
- Whether any complementary interview lenses were applied (landscape mapping, choice cascade, value innovation) — if so, note which ones for the interview lens audit in Step 3
- The stated scope and timeframe
- The intended audience
- Any assumptions listed (or conspicuously absent)
Step 3 — Evaluate against the seven dimensions
Work through each dimension from references/review-dimensions.md. For each:
1. Assess — What does the document do well or poorly on this dimension? 2. Find evidence — Quote or cite specific passages. Don't make claims you can't point to. 3. Rate — Assign a rating (Strong / Adequate / Weak / Missing) only after step 2 satisfies Hard gates → Ratings (anchor on record before the label). 4. Recommend — If Weak or Missing, what specifically should the author do? Not "improve the diagnosis" — rather, "the diagnosis names 'market shift' as the challenge but doesn't specify which shift or why it matters for this company specifically. A stronger version would identify the one or two structural changes that create the opening or threat."
The seven dimensions (summarized here, detailed in the reference):
| # | Dimension | Core question |
|---|---|---|
| 1 | Diagnosis quality | Does it name the actual challenge with enough specificity to be wrong? |
| 2 | Guiding policy strength | Does it make a real choice that rules things out? |
| 3 | Action coherence | Do the actions reinforce each other and carry out the policy? |
| 4 | Kernel chain | Does the logical chain from diagnosis -> policy -> actions hold? |
| 5 | Bad-strategy patterns | Are any of the four Rumelt hallmarks (plus one additional anti-pattern) present? |
| 6 | Assumption exposure | Are load-bearing assumptions identified and testable? |
| 7 | Specificity and falsifiability | Could you tell in 12 months whether this strategy worked or failed? |
Step 3b — Interview lens audit (when interview lenses appear in the document)
If the strategy document or its companion notes show evidence that landscape mapping (Wardley), strategic choice cascade (Playing to Win), or value innovation (Blue Ocean) lenses were applied during the interview, audit whether those lens findings actually made it into the strategy. Load references/interview-lens-audit.md for the detailed audit criteria for each lens.
The audit checks two things for each lens: 1. Fidelity — Did the lens's specific findings survive into the draft, or were they dropped or softened? 2. Impact — Did the lens findings actually sharpen the kernel (diagnosis, guiding policy, coherent actions), or were they acknowledged but left decorative?
Record findings under the relevant dimensions (typically Dim 1 for landscape, Dim 2-3 for cascade, Dim 1-4 for value innovation). Also flag any interview lens findings from the notes that were dropped or softened in the draft — these often represent the sharpest thinking that got smoothed away during document polishing.
Note: The four review lenses (7S, Scorecard, Five Forces, Hoshin Kanri) and the three interview lenses serve different purposes. Review lenses are tools the reviewer applies to find gaps. Interview lenses produced findings during the strategy-building conversation — this audit checks whether those findings survived. Don't substitute one for the other: Five Forces can pressure-test competitive claims independently, but it cannot validate whether a Wardley mapping exercise was faithfully represented in the draft.
Step 4 — Cross-check with notes (if available)
If strategy-notes.md or equivalent reasoning notes are available, cross-reference:
- Open questions: Are any of these actually show-stoppers that should block the draft from being finalized?
- Assumptions: Does the draft rest on assumptions the notes flagged as unverified? If so, how severe is the exposure?
- Bad-strategy patterns caught during interview: Were they actually resolved in the final draft, or did they creep back in softer language?
- Thinking evolution: Did the draft preserve the sharpest version of the diagnosis, or did it soften during writing?
- Lens findings: If landscape mapping, choice cascade, or value innovation analysis was done, did those findings actually make it into the draft?
Step 5 — Produce the review
Before writing files, check the user's intent. If they asked for quick feedback, a chat-only take, or to "poke holes in this" conversationally, deliver the findings inline — lead with Critical Findings, then the top failure path, then recommendations. Offer to write the full review file afterward. If the user asked for a formal review or didn't specify, confirm the output path before writing.
Write strategy-review.md following references/review-template.md. The structure:
1. Review summary — 3-4 sentences: what this strategy is trying to do, the overall assessment, and the one thing that most needs attention. 2. Strengths — What works well and why. Be specific. Start here so the author knows what to protect. 3. Critical findings — The 2-4 things that most undermine the strategy's integrity, ordered by severity. These are the items that should be fixed before the strategy is shared or acted on. Lead with these — they are the highest-value part of the review. 4. Dimension ratings — The seven dimensions with ratings, evidence, and recommendations. These provide the supporting detail behind the critical findings. 5. Failure path analysis — How this strategy could fail in the real world. See "Thinking about failure paths" below. 6. Assumption risk map — Load-bearing assumptions ranked by impact x uncertainty. Which ones, if wrong, would break the strategy entirely? 7. Lens findings — If any review lenses were applied, what they revealed that the core dimensions alone didn't catch. Omit this section if no lenses were triggered. 8. Recommended next steps — Concrete actions to strengthen the document, ordered by impact.
Compose from artifacts, not memory. If durable review state exists (see below), use those files as the primary source for the final review rather than reconstructing from the conversation. Update review-composition.md with the overall assessment and structure decisions, then compose strategy-review.md from the artifacts. If the agent supports subagents, dispatch one for document assembly — a fresh context reading structured files produces more accurate reviews; otherwise re-read the artifact files directly and compose them yourself — identical output.
After writing the file, give a brief chat summary: overall assessment in one sentence, the single most important finding, and the recommended next action.
Thinking about failure paths
This is where the review delivers the most value that the author can't easily get elsewhere. Most people reviewing a strategy will tell you what's wrong with the document. Failure path analysis tells you what could go wrong in the real world because of what's in (or missing from) the document.
Strategies rarely fail through dramatic, visible crisis. They fail through predictable, quiet patterns:
The capability gap: The load-bearing action requires a capability the organization doesn't have and underestimates the cost of building. "Ship a native mobile app by Q3" sounds like an action item; if the team has never built native mobile, it's a six-month research project dressed up as a deliverable.
The slow drift: Nothing forces adherence to the guiding policy's exclusions. Quarter by quarter, exceptions creep in. "We said no enterprise, but this one deal is really big." Within a year, the strategy has been silently reversed by a thousand small yeses.
The unmodeled response: The strategy assumes the competitive environment is static. But the competitor reads the same market signals. If the strategy's diagnosis is correct, others will reach similar conclusions — what happens when they respond?
The political veto: The strategy requires stopping something that a powerful stakeholder owns. The strategy document doesn't mention this because the author assumes buy-in will materialize. It usually doesn't.
The assumption cascade: One key assumption turns out to be wrong, and because the actions are tightly coupled, the failure propagates. The market doesn't grow as expected, so the unit economics don't work, so the funding case falls apart, so the capability never gets built.
The execution bottleneck: Multiple actions depend on the same scarce resource (a key team, a specific leader, a budget line) but the strategy treats them as independent. When the bottleneck binds, the author must choose which actions to delay — and the strategy didn't provide guidance on that priority call.
The measurement void: No leading indicators were defined, so the organization can't tell whether the strategy is working until lagging indicators arrive (revenue, market share) — by which time it's too late to course-correct.
For each failure path you identify, describe: the scenario in concrete terms, the likelihood, what the strategy currently does to mitigate it (often nothing), and what the author could do to reduce the risk. These are the findings authors most often say changed their thinking — because they expose risks the author knew existed but hadn't made explicit.
Posture and calibration
Be honest, not hostile
The author spent real thinking effort on this. Acknowledge what works — genuinely, not as a softener before the bad news. But don't pull punches on structural problems. A strategy that goes to the board with a flawed diagnosis will do more damage than a reviewer who was too direct.
Frame findings as structural observations, not personal criticism: "The diagnosis addresses market dynamics but doesn't identify why this company is specifically exposed" — not "you didn't think hard enough about the diagnosis."
Calibrate to the document's maturity
- Early draft / working document: Focus on the big structural issues. Don't nitpick prose or completeness — the author knows it's rough. Emphasize what's promising and what needs the most work.
- Near-final / about to be shared: Higher bar. Flag everything that could undermine credibility with the audience. Check that the "At a glance" section accurately represents the full document. Verify that claimed exclusions in the guiding policy aren't quietly reintroduced in the actions.
- Post-hoc / strategy already in execution: Focus on assumption validation. Which assumptions can now be checked against reality? Flag any that have already been falsified by events.
Don't rewrite the strategy
The review identifies problems and recommends fixes. It does not produce an alternative strategy. If the diagnosis is fundamentally wrong, say so and explain why — but the user needs to rethink it themselves (ideally using the strategy-interview skill). A reviewer who rewrites the strategy is doing the author's thinking for them, which means the author won't own it.
Exception: if the document has a missing element (no diagnosis at all, no stated exclusions), offer a concrete example of what a good version might look like, clearly marked as illustrative, to help the author see the gap.
Respect different frameworks
If the document uses OKRs, V2MOM, SWOT, Porter's Five Forces, or any other framework — evaluate the thinking, not the format. The kernel question is always the same: is there a diagnosis, a directional choice, and coherent actions? These may appear under different labels. Map concepts, don't force vocabulary.
Flag genuine structural gaps ("this V2MOM has methods and obstacles but no diagnosis of which obstacle matters most") rather than framework translation ("you should use the kernel format instead").
Source discipline
When the document makes claims about markets, competitors, or trends:
- Note which claims are sourced and which are asserted without evidence.
- Flag market claims that feel like conventional wisdom rather than verified data.
- If the review depends on external facts you can't verify, say so: "The diagnosis rests on the claim that [X] — I can't verify this, but if it's wrong, the entire strategy changes."
Variant: reviewing a strategy-interview in progress
If the user is mid-interview (they have strategy-notes.md but no strategy-draft.md, or the draft is marked [PROVISIONAL]):
1. Review what exists so far — the notes capture the thinking even before the draft crystallizes. 2. Focus on the strongest and weakest elements of the emerging kernel. 3. Suggest questions the user should pressure-test before finalizing. 4. Don't produce a full review — produce a "mid-interview check" with the 3-5 most important observations.
Variant: comparative review
If the user provides two strategy documents (e.g., competing proposals, before/after versions, or strategies from different teams):
1. Review each independently first using the seven dimensions. 2. Then compare: where do they agree on the diagnosis? Where do they diverge? Is one more specific, more coherent, or more falsifiable? 3. Don't declare a winner unless the structural quality is clearly different. Two strategies can be equally rigorous and still disagree — that's a judgment call for the decision-maker, not the reviewer.
Durable review state
Long or complex reviews lose fidelity — evidence citations drift, dimension assessments reference passages from memory instead of the document, and the final review sounds more certain than the source material supports. For reviews that warrant it, maintain working state in a .beagle/ directory.
When to start
Create the directory when any of these appear:
- The input document exceeds roughly 2,000 words or spans multiple files.
- The review is comparative (two or more strategy documents).
- The strategy is near-final or about to be shared with stakeholders.
- Two or more review lenses are activated.
- The document includes sourced market claims that need separate tracking.
Directory location
- For interview-produced artifacts:
.beagle/strategy/<subject-slug>/reviews/<date-or-slug>/— nests the review alongside the interview state. - For standalone reviews:
.beagle/strategy-review/<subject-slug>/reviews/<date-or-doc-slug>/— supports re-reviewing the same subject or reviewing v1/v2 documents without mixing evidence.
Working files
| File | Purpose | Created when |
|---|---|---|
review-state.md | Review ledger — document metadata, lenses triggered, current step | Always, once directory exists |
source-evidence.md | Document quotes and claims with provenance tags | When the document is read (Step 2) |
kernel-extraction.md | Extracted kernel elements with evidence | After Step 2 |
dimension-ratings.md | Per-dimension assessment, evidence, and ratings | During Step 3 |
assumption-risk-map.md | Load-bearing assumptions with impact × uncertainty | During Step 3/4 |
failure-paths.md | Failure scenarios with likelihood and mitigation | After dimension evaluation |
lens-notes.md | Lens-specific findings | When a lens is used |
review-composition.md | Pre-composition outline for the final review | Before Step 5 |
Evidence tagging (source-evidence.md)
Tag each entry to prevent the final review from sounding more certain than the source material supports:
- `doc quote` — direct quote from the strategy document, with section reference.
- `reviewer inference` — derived from the document, not directly stated.
- `unverified assumption` — claim the strategy depends on, not confirmed in the document.
- `source-backed` — claim in the document that cites a source.
- `notes cross-ref` — finding from strategy-notes.md or interview artifacts.
- [doc quote] "We will focus exclusively on mid-market SaaS" (Guiding Policy, p.3)
- [reviewer inference] Mid-market focus implies exiting current enterprise contracts. (Dim 2)
- [unverified assumption] TAM estimate of $2B appears unsourced. (Dim 6)
- [source-backed] "Mobile inspection usage grew 68% YoY (Gartner 2025)" (Diagnosis, p.1)
- [notes cross-ref] Interview notes flagged enterprise exit as contested. (Notes)Review state ledger (review-state.md)
document: [filename(s) being reviewed]
subject: [what the strategy is for]
maturity: [early-draft / near-final / post-hoc]
audience: [who reads the strategy]
current_step: [1-5]
lenses_triggered: [7S, scorecard, five-forces, hoshin-kanri, or none]
kernel_found: [yes / partial / no]
diagnosis_summary: [one sentence]
policy_summary: [one sentence]
critical_findings_count: [number]
top_risk: [one sentence]Optional: judge artifact mode
Normal reviews produce prose (strategy-review.md). When the user explicitly requests machine-readable output for evaluation pipelines, RL loops, or LLM-as-judge workflows, produce an additional structured JSON artifact alongside the prose.
When to activate
Judge artifact mode activates only when the user's language signals it. Look for phrases like:
- "machine-readable," "structured output," "JSON report"
- "LLM judge," "LLM-as-judge," "judge loop," "evaluation artifact"
- "RL loop," "reward signal," "evaluation pipeline"
- "produce the JSON," "judge artifact," "structured review"
Do not activate for normal reviews, quick takes, or chat-only feedback. If in doubt, ask: "Want me to also produce a machine-readable JSON artifact for evaluation pipelines, or just the prose review?"
What it produces
In judge artifact mode, produce both files unless the user asks for JSON-only:
1. `strategy-review.md` — the standard prose review (written first) 2. `strategy-review.json` — structured JSON following references/judge-artifact-schema.md
Write the prose review first, then derive the JSON from it. This ensures scores emerge from careful analysis, not from filling a schema. Every dimension label, critical finding, and evidence quote in the JSON must be consistent with the prose.
Production rules
1. Scores are projections, not new judgments. The dimension ratings in JSON map directly from the prose labels: Strong=4, Adequate=3, Weak=2, Missing=1. Do not invent intermediate scores. 2. Every score must cite evidence. At least one evidence entry per dimension, each tagged with provenance (doc_quote, reviewer_inference, unverified_assumption, source_backed, or notes_cross_ref). 3. Distinguish document facts from reviewer inferences. A downstream consumer filtering by provenance should be able to separate what the document says from what the reviewer concluded. This is the single most important quality signal for evaluation pipelines. 4. All seven dimensions and all five bad-strategy patterns are always present. Use detected: false for absent patterns. This gives consumers stable field counts. 5. Optional signals when produced in prose. The schema also accepts an optional top-level strengths array, a top-level recommended_next_steps array, a review_metadata.review_id identifier, and a notes_cross_reference object (with patterns_that_crept_back and thinking_sharpened_then_softened sub-arrays). When the prose review produces the matching sections — and, for notes_cross_reference, when notes were available — emit the corresponding JSON fields. Otherwise omit them rather than emitting empty stubs. 6. Blocking findings gate the reward signal. A strategy passes (reward_signal.pass: true) only if the aggregate score is at least Adequate AND no findings are marked blocking. 7. Validate before emitting. JSON must parse, required fields must exist, scores must match labels, weighted total must be arithmetically correct. See the validation checklist in the schema reference.
Aggregate scoring
The aggregate is a weighted composite of the seven dimension scores:
| Dimension | Weight |
|---|---|
| Diagnosis quality | 20% |
| Guiding policy strength | 20% |
| Action coherence | 15% |
| Kernel chain integrity | 15% |
| Bad-strategy patterns | 10% |
| Assumption exposure | 10% |
| Specificity / falsifiability | 10% |
weighted_total = Σ(score × weight / 100). Range: 1.0–4.0. The aggregate label follows: Strong (≥3.5), Adequate (≥2.5), Weak (≥1.5), Missing (<1.5).
The reward signal normalizes to 0.0–1.0: (weighted_total - 1.0) / 3.0.
Source of truth: The canonical weight definitions, threshold table, and validation
rules live in references/judge-artifact-schema.md. If these values are updated,update both locations.
Reference files
references/review-dimensions.md— The seven evaluation dimensions with detailed criteria, examples, and common failure patterns.references/review-lenses.md— Four complementary review lenses (7S, Scorecard, Five Forces, Hoshin Kanri) with triggers, questions, and common gaps.references/interview-lens-audit.md— Audit criteria for checking whether interview lens findings (Wardley, Playing to Win, Blue Ocean) survived into the draft.references/review-template.md— Exact structure of thestrategy-review.mdoutput file.references/judge-artifact-schema.md— JSON schema for machine-readable judge artifacts. Defines field structure, score mapping, evidence provenance rules, validation checklist, and relationship to prose output.references/pressure-tests.md— Expected behaviors for common review scenarios. For skill validation.
Interview Lens Audit
Audit criteria for checking whether interview lens findings survived into the strategy draft. Use this during Step 3b when the strategy document or companion notes show evidence that landscape mapping, strategic choice cascade, or value innovation lenses were applied during the strategy-interview skill.
The audit is about fidelity and impact — did the lens findings make it into the draft, and did they actually sharpen the kernel? A lens finding that appears in the notes but vanishes from the draft is a red flag. A lens finding that appears in the draft but didn't change anything is decorative.
---
1. Landscape Mapping (Wardley)
What to look for in the notes
Wardley-informed interview notes typically contain:
- A value chain or dependency structure (user need at the top, components underneath)
- Evolution stage assessments for key components (genesis, custom-built, product, commodity)
- Tension points: building custom what's becoming commodity, treating genesis as procurement, competitor further along the evolution curve
- Inertia patterns: attachment to a component whose evolution stage has shifted
- Value-capture dynamics: which layer of the chain captures margin, and whether that's shifting
Audit checklist
Check each item against both the notes and the draft. A "present in notes, absent in draft" finding is the most important signal.
| Criterion | What to check | Dimension |
|---|---|---|
| User need anchoring | Does the diagnosis start from a clearly stated user need, or does it start from the company's internal perspective? Wardley mapping anchors everything to the user need at the top of the chain. | Dim 1 |
| Value-chain / dependency structure | Does the diagnosis identify which components the strategy depends on and how they relate? Not necessarily a formal map — but awareness of the dependency chain. | Dim 1 |
| Evolution stage awareness | Does the diagnosis distinguish between components at different evolution stages? Look for language about commoditization, maturity, or "table stakes" vs. novel/uncertain capabilities. | Dim 1 |
| Control and dependency points | Does the strategy identify which components it controls vs. depends on? A strategy built on a component controlled by someone else carries risk the document should acknowledge. | Dim 6 |
| Value-capture or inertia insights | Did the interview surface where margin sits in the chain, or where organizational inertia is masking a shift? Does the draft reflect these insights, or did it revert to conventional framing? | Dim 1, Dim 6 |
| Diagnosis sharpening | Did the landscape analysis actually change or sharpen the diagnosis — or was it acknowledged but left aside? A good test: does the diagnosis contain a claim that could only come from mapping the landscape (e.g., "we're building custom what has become commodity"), or does it read the same as it would without the mapping? | Dim 4 |
Common drop patterns
These are the ways Wardley findings most often get lost between notes and draft:
- Abstraction wash: The notes contain specific evolution-stage findings ("our data pipeline is custom-built but three vendors now sell this as a product") but the draft abstracts them into generic competitive language ("we face increasing competitive pressure").
- Inertia silence: The notes identify organizational inertia around a commoditizing capability, but the draft treats the current approach as a strength rather than a liability.
- Missing dependency risk: The notes flag a critical dependency on a component controlled by another party, but the draft's assumption exposure doesn't mention it.
- Value-chain flattening: The notes describe a multi-layer value chain with different dynamics at each layer, but the draft collapses it into a single "our market" description.
---
2. Strategic Choice Cascade (Playing to Win)
What to look for in the notes
Cascade-informed interview notes typically contain:
- A winning aspiration — concrete picture of what winning looks like
- Where-to-play choices — specific segments, geographies, or value-chain positions chosen (and excluded)
- How-to-win mechanism — the structural advantage, not just "better execution"
- Capability assessment — 3-5 reinforcing capabilities required, with gap identification
- Management systems — processes, metrics, and feedback loops to keep the strategy on track
Audit checklist
| Criterion | What to check | Dimension |
|---|---|---|
| Winning aspiration clarity | Does the draft contain a concrete picture of success that two people would agree on, or has it softened into a generic vision statement? | Dim 7 |
| Where-to-play specificity | Does the guiding policy name specific segments and exclusions? If the cascade surfaced a narrow playing field, check whether the draft expanded it back to "everyone." | Dim 2 |
| How-to-win mechanism | Does the guiding policy name a structural advantage — not just "superior product" but the specific asymmetry? If the cascade identified why competitors can't copy the advantage, does that reasoning appear? | Dim 2 |
| Capability gaps in assumptions | Did capability gaps identified during the cascade appear in the assumption exposure or as caveats in the actions? Or were they silently assumed away? | Dim 6 |
| Management systems in actions | Did the cascade's management systems findings — measurement, reinforcement, drift prevention — make it into the coherent actions? This is the most commonly dropped cascade element. | Dim 3 |
| Policy/action sharpening | Did the cascade choices actually tighten the guiding policy and actions, or were they acknowledged in passing? Test: remove the cascade language — does the policy read differently? If not, the cascade was decorative. | Dim 4 |
Common drop patterns
- Aspiration inflation: The cascade produced a specific, bounded winning aspiration, but the draft inflated it into a broad vision statement.
- Exclusion erosion: The where-to-play analysis excluded specific segments, but the draft hedges ("primarily mid-market, but we'll also serve enterprise when opportunities arise").
- Capability optimism: The cascade identified capability gaps, but the draft treats them as tasks rather than risks — "hire mobile team" instead of acknowledging a fundamental capability the org hasn't built.
- Systems omission: The most common drop. Management systems are almost always the weakest part of the cascade, and the first thing cut from the draft.
---
3. Value Innovation (Blue Ocean)
What to look for in the notes
Blue Ocean-informed interview notes typically contain:
- A strategy canvas or convergence analysis — factors the industry competes on, with all players roughly similar
- ERRC moves — specific factors to eliminate, reduce, raise, or create
- Noncustomer tier analysis — who isn't buying and why
- A diagnosis reframe from "how to beat competitors" to "the competitive arena itself is the problem"
Audit checklist
| Criterion | What to check | Dimension |
|---|---|---|
| Strategy canvas / convergence recognition | Does the diagnosis acknowledge competitive convergence — that all players compete on the same factors? Or did it revert to "we're behind on features"? | Dim 1 |
| ERRC moves in actions | Do the coherent actions reflect eliminate/reduce/raise/create moves — deliberately diverging from competitors? Or do they default to matching and incrementing? | Dim 3 |
| Noncustomer awareness | If noncustomer tiers were identified, does the strategy target them, or did the draft narrow back to existing customers? | Dim 2 |
| Diagnosis reframe | Did the diagnosis shift from competitive positioning ("we're losing share") to competitive convergence ("the category has converged and further investment in these factors produces declining returns")? This is the most important Blue Ocean contribution to the kernel. | Dim 1 |
| Value-cost tradeoff | Does the guiding policy break the value-cost tradeoff (simultaneously reducing cost in some areas and increasing value in others), or does it stay within the conventional differentiation-vs-cost frame? | Dim 2, Dim 4 |
| Divergence from competitor matching | Read the coherent actions as a whole: are they about catching up to competitors, or about going somewhere competitors aren't? If the Blue Ocean analysis surfaced a divergent path but the actions converge back, flag it. | Dim 3, Dim 4 |
Common drop patterns
- Convergence denial: The notes document clear competitive convergence, but the draft frames the problem as "we need to execute better on the same dimensions" — reverting to a red ocean diagnosis.
- ERRC flattening: The notes contain specific eliminate/reduce/raise/create moves, but the draft turns them into conventional feature prioritization ("we'll focus on X and Y") without the deliberate elimination and reduction that makes the strategy distinctive.
- Noncustomer retreat: The notes identify a promising noncustomer tier, but the draft's playing field narrows back to the existing market — often because the existing market feels safer or more measurable.
- Incremental creep: The actions start divergent but accumulate "and also" items that pull the strategy back toward competitor matching. Check whether actions added after the Blue Ocean analysis dilute the ERRC moves.
---
Reporting interview lens audit findings
Record each finding under the relevant dimension in the dimension ratings. In the Lens Findings section of the review output, use this format for the interview lens audit subsection:
**Interview lens audit:**
- **[Lens name] — [Reflected / Partially reflected / Missing]**: [What survived, what was dropped or softened, and what was lost. If partially reflected, name the specific elements that made it and the ones that didn't.]The most valuable finding is often the gap between the notes and the draft — the sharpest strategic thinking from the interview that got polished away during document writing. Name it specifically so the author can recover it.
Judge Artifact Schema
Machine-readable JSON output for strategy reviews. Produced only when judge artifact mode is explicitly requested — never emitted during normal reviews.
The schema bridges the prose review (strategy-review.md) and automated evaluation pipelines. Every numeric score maps directly to a prose label the reviewer already assigned. No new judgments are invented for the JSON — it's a structured projection of the same review, not a separate evaluation.
---
Output file
Filename: strategy-review.json Location: Same directory as strategy-review.md. If durable review state is active, also write to the review state directory.
When the user requests both prose and JSON (the default in judge mode), produce strategy-review.md first, then strategy-review.json. The JSON must be consistent with the prose — if a dimension is rated "Weak" in the prose, the JSON score must be 2.
---
Score mapping
Strategy-review uses four rating levels. Map them to integers:
| Label | Score | Description |
|---|---|---|
| Strong | 4 | Element is specific, testable, and structurally sound |
| Adequate | 3 | Element is present and functional but has notable gaps |
| Weak | 2 | Element has significant structural problems |
| Missing | 1 | Element is absent or so vague it provides no guidance |
This is a 1–4 scale, not 1–5. The strategy-review skill has four levels; the JSON reflects them faithfully rather than inventing intermediate granularity.
---
Schema definition
Note: The example below is a partial illustrative fragment. It shows two of seven requireddimension_ratings, one of the minimum twofailure_paths, and one of the required 2–4critical_findingsto demonstrate the structure without duplicating the pattern for every entry. A validstrategy-review.jsonmust include all seven dimensions, at least two failure paths, and 2–4 critical findings per the field reference below. The example also illustrates the optionalreview_id,strengths,notes_cross_reference, andrecommended_next_stepsfields — these may be omitted when not applicable.
{
"schema_version": "1.0",
"evaluator": {
"skill": "strategy-review",
"version": "2.11.0"
},
"review_metadata": {
"review_id": "fieldkit-mobile-2026-04-10-near-final",
"reviewed_at": "2026-04-10T14:30:00Z",
"documents": [
{
"name": "strategy-draft.md",
"path": "relative/path/or/null",
"sha256": "hex digest or null",
"word_count": 1850
}
],
"subject": "Mobile-first inspection workflows for mid-market",
"maturity": "near-final",
"audience": "Board of directors",
"timeframe": "18 months"
},
"kernel_extraction": {
"diagnosis": {
"found": true,
"summary": "Growth stalled because competitors launched mobile-first while Fieldkit remained desktop-only.",
"evidence": [
{
"text": "68% of inspections happen on-site with a phone, but our mobile experience is a responsive web wrapper that drops offline.",
"provenance": "doc_quote",
"location": "Diagnosis, paragraph 2"
}
]
},
"guiding_policy": {
"found": true,
"summary": "Dominate mobile-first inspection workflows for mid-market; stop pursuing enterprise until core product is defensible.",
"exclusions": [
"Enterprise sales motion",
"SOC 2 / SSO / audit trail features this year",
"Matching BuildOps on feature breadth"
],
"evidence": [
{
"text": "Won't build enterprise compliance features this year.",
"provenance": "doc_quote",
"location": "Guiding Policy"
}
]
},
"coherent_actions": {
"found": true,
"count": 4,
"actions": [
"Ship native mobile app with full offline sync by Q3",
"Kill enterprise sales motion — reassign AEs to mid-market",
"Build HappyCo and Yardi integrations",
"Launch reliability guarantee"
],
"evidence": [
{
"text": "Kill enterprise sales motion — reassign two enterprise AEs to mid-market, cancel SOC 2 engagement ($180K → mobile engineering)",
"provenance": "doc_quote",
"location": "Coherent Actions, item 2"
}
]
},
"chain_paragraph": "Growth stalled because competitors launched mobile-first while Fieldkit remained desktop-only, and the company drifted into enterprise to compensate. Therefore, dominate mobile-first inspection workflows for mid-market and stop pursuing enterprise until the core product is defensible. Which means: ship native mobile with offline sync, kill enterprise sales, build mid-market integrations, and launch a reliability guarantee."
},
"dimension_ratings": [
{
"dimension": 1,
"name": "diagnosis_quality",
"score": 4,
"label": "Strong",
"confidence": "high",
"assessment": "Names a specific structural mistake with concrete evidence. The 68% mobile statistic and the drift-to-enterprise pattern are falsifiable claims.",
"evidence": [
{
"text": "Treating mobile as feature request, not delivery surface.",
"provenance": "doc_quote",
"location": "Diagnosis"
}
],
"recommendation": null
},
{
"dimension": 2,
"name": "guiding_policy_strength",
"score": 4,
"label": "Strong",
"confidence": "high",
"assessment": "Makes a painful cut (enterprise) to concentrate force where the org has an edge. Exclusions are explicit.",
"evidence": [
{
"text": "Won't build enterprise compliance features this year. Won't match BuildOps on breadth.",
"provenance": "doc_quote",
"location": "Guiding Policy"
}
],
"recommendation": null
}
],
"bad_strategy_patterns": [
{
"pattern": "fluff",
"detected": false,
"severity": null,
"evidence": [],
"resolution": null
},
{
"pattern": "failure_to_face_challenge",
"detected": false,
"severity": null,
"evidence": [],
"resolution": null
},
{
"pattern": "goals_as_strategy",
"detected": false,
"severity": null,
"evidence": [],
"resolution": null
},
{
"pattern": "bad_objectives",
"detected": false,
"severity": null,
"evidence": [],
"resolution": null
},
{
"pattern": "strategy_by_analogy",
"detected": false,
"severity": null,
"evidence": [],
"resolution": null
}
],
"assumption_risk_map": [
{
"assumption": "Team can ship native mobile with offline sync by Q3",
"stated_in_doc": false,
"impact_if_wrong": "high",
"uncertainty": "high",
"risk_level": "critical",
"provenance": "reviewer_inference",
"consequence": "Load-bearing action fails. Reliability guarantee becomes liability. Integrations built on unstable foundation."
}
],
"failure_paths": [
{
"name": "Capability gap stalls native mobile",
"pattern": "capability_gap",
"scenario": "Team has never built native mobile. Q3 deadline assumes execution speed the org hasn't demonstrated. Six-month research project dressed as a deliverable.",
"likelihood": "medium",
"current_mitigation": "Budget reallocation from SOC 2 engagement",
"suggested_mitigation": "Prototype sprint in month 1. If offline sync isn't working by week 6, the Q3 date is fiction — trigger contingency plan."
}
],
"review_lenses_applied": [
{
"lens": "balanced_scorecard",
"trigger": "Success criteria are financial-only (ARR growth target)",
"findings": [
"No leading indicators from customer or internal-process perspectives. Strategy has a 6-month blind spot between cause and detection."
]
}
],
"interview_lens_audit": [
{
"lens": "landscape_mapping",
"assessment": "partially_reflected",
"survived": [
"Evolution stage awareness for mobile delivery surface"
],
"dropped": [
"Notes identified data pipeline as custom-built when three vendors now offer it as product — draft omits this"
],
"impact_on_kernel": "Diagnosis sharpened by mobile evolution insight but missed the infrastructure commoditization finding",
"evidence": [
{
"text": "Our data pipeline is custom-built but three vendors now sell this as a product",
"provenance": "notes_cross_ref",
"location": "lens-notes.md, landscape mapping section"
},
{
"text": "We face competitive pressure in infrastructure",
"provenance": "doc_quote",
"location": "Diagnosis, paragraph 3"
}
]
}
],
"strengths": [
{
"title": "Painful exclusions are explicit",
"description": "Names what it won't do (enterprise, SOC 2, BuildOps feature parity) — exclusions concentrate force where the team has an edge.",
"evidence": [
{
"text": "Won't build enterprise compliance features this year. Won't match BuildOps on breadth.",
"provenance": "doc_quote",
"location": "Guiding Policy"
}
]
},
{
"title": "Action set is coherent and mutually reinforcing",
"description": "Actions align to the same policy choice and reinforce the mid-market mobile focus.",
"evidence": [
{
"text": "Ship native mobile app with full offline sync by Q3",
"provenance": "doc_quote",
"location": "Coherent Actions, item 1"
}
]
}
],
"critical_findings": [
{
"title": "Native mobile capability assumed but unverified",
"severity": "critical",
"description": "The load-bearing action (native mobile app with offline sync by Q3) requires a capability the team has never demonstrated.",
"impact": "If this action slips, the reliability guarantee becomes a liability and the enterprise exit loses its rationale.",
"evidence": [
{
"text": "Ship native mobile app with full offline sync by Q3",
"provenance": "doc_quote",
"location": "Coherent Actions, item 1"
}
],
"recommendation": "Add a capability validation milestone in month 1. Define what 'on track' looks like at week 6."
}
],
"blocking_findings": [
"Native mobile capability assumed but unverified"
],
"unresolved_questions": [
"Has the team built native mobile before, or is this a first attempt?",
"What happens to existing enterprise customers during the transition?"
],
"notes_cross_reference": {
"patterns_that_crept_back": [
{
"pattern": "fluff",
"description": "Notes flagged 'leverage synergies' as fluff during the interview; the draft replaced it with 'capture cross-product value' — softer wording, same emptiness.",
"notes_quote": "Drop 'leverage synergies' — meaningless.",
"draft_quote": "Capture cross-product value across mobile and web."
}
],
"thinking_sharpened_then_softened": [
{
"topic": "diagnosis",
"notes_version": "Drift to enterprise was an admission we couldn't win mid-market on product, not a strategic choice.",
"draft_version": "Pursued enterprise opportunities in parallel with mid-market growth.",
"what_was_lost": "Notes named the structural failure (couldn't win mid-market). Draft reframes it as a parallel opportunity, masking the diagnosis."
}
]
},
"recommended_next_steps": [
{
"action": "Add a capability validation milestone in month 1 with an explicit week-6 trigger.",
"rationale": "Native mobile is the load-bearing action. Without an early check, Q3 failure is invisible until it's too late."
},
{
"action": "Define the metric that proves mid-market dominance.",
"rationale": "Strategy passes structural review but lacks a specificity hook the team can point to in 12 months."
}
],
"aggregate_score": {
"weighted_total": 3.45,
"weights": {
"diagnosis_quality": 20,
"guiding_policy_strength": 20,
"action_coherence": 15,
"kernel_chain_integrity": 15,
"bad_strategy_patterns": 10,
"assumption_exposure": 10,
"specificity_falsifiability": 10
},
"label": "Adequate",
"confidence": "high"
},
"reward_signal": {
"pass": false,
"score": 0.82,
"rationale": "One critical blocking finding (unverified capability assumption on load-bearing action). Strategy is structurally sound but not ready to act on without resolving the capability question."
}
}---
Field reference
Top-level fields
| Field | Type | Required | Description |
|---|---|---|---|
schema_version | string | yes | Schema version. Currently "1.0". |
evaluator | object | yes | Skill identifier and version. |
review_metadata | object | yes | Document metadata and review context. |
kernel_extraction | object | yes | Extracted kernel elements with evidence. |
dimension_ratings | array[7] | yes | All seven dimensions, each with score and evidence. |
bad_strategy_patterns | array[5] | yes | All five patterns, each with detected flag. Always include all five — detected: false for absent patterns. |
assumption_risk_map | array | yes | Load-bearing assumptions. May be empty if none found (unusual). |
failure_paths | array | yes | Failure scenarios. Minimum 2, typically 3–5. |
review_lenses_applied | array | no | Omit if no review lenses triggered. |
interview_lens_audit | array | no | Omit if no interview lenses were used during the strategy-interview. |
notes_cross_reference | object | no | Structured signals from a notes-vs-draft cross-read. Omit if no strategy-notes.md or interview durable state was available, or if both sub-arrays would be empty. |
strengths | array | no | Specific strengths the author should protect. 2–4 entries when emitted. Mirrors the prose "What Works" section. |
critical_findings | array | yes | The 2–4 highest-severity findings. |
blocking_findings | array | yes | Titles of findings that should block the strategy from being finalized. May be empty. |
unresolved_questions | array | yes | Questions the review couldn't answer. May be empty. |
recommended_next_steps | array | no | 2–4 priority actions for post-review follow-up. Mirrors the prose "Recommended Next Steps" section. Array order encodes priority. Omit if not produced. |
aggregate_score | object | yes | Weighted composite score. |
reward_signal | object | yes | Pass/fail determination for evaluation loops. |
Review metadata fields
| Field | Type | Required | Description |
|---|---|---|---|
review_id | string | no | Stable identifier for this review run. Used by pipelines to deduplicate when the same document is reviewed multiple times. Format is caller's choice — UUIDv4 (e.g., 550e8400-e29b-41d4-a716-446655440000) or a slug like <doc-slug>-<YYYY-MM-DD>-<maturity> are both valid. Must be unique per review run, not per document. |
reviewed_at | string (ISO 8601) | yes | Timestamp of the review. |
documents | array | yes | Documents reviewed, each with name, path, sha256, word_count. |
subject | string | yes | What the strategy is about. |
maturity | enum | yes | Document maturity stage. One of: early-draft, near-final, post-hoc. |
audience | string | yes | Intended audience for the strategy. |
timeframe | string | no | Strategy timeframe if stated in the document. |
Evidence objects
Evidence entries appear throughout the schema. Every evidence object has:
| Field | Type | Required | Description |
|---|---|---|---|
text | string | yes | The quoted passage or stated finding. |
provenance | enum | yes | One of: doc_quote, reviewer_inference, unverified_assumption, source_backed, notes_cross_ref. |
location | string | no | Section or page reference in the source document. |
Provenance rules — these match the evidence tagging in source-evidence.md:
- `doc_quote`: Direct quote from the strategy document. Must be verifiable against the source.
- `reviewer_inference`: Derived from the document but not directly stated. The reviewer connected dots the author didn't.
- `unverified_assumption`: Claim the strategy depends on that isn't confirmed in the document or notes.
- `source_backed`: Claim in the document that cites an external source.
- `notes_cross_ref`: Finding from strategy-notes.md or interview durable state artifacts.
Notation mapping: The durable review state files (source-evidence.md) use a slightly different format for the same tags: doc quote, reviewer inference, unverified assumption, source-backed, notes cross-ref. The JSON enum values use underscores (doc_quote, source_backed, notes_cross_ref) per JSON convention. The semantics are identical.
Dimension rating objects
| Field | Type | Required | Description |
|---|---|---|---|
dimension | integer | yes | 1–7, matching the review dimensions. |
name | string | yes | Snake_case dimension name (e.g., diagnosis_quality). |
score | integer | yes | 1–4. See score mapping above. |
label | string | yes | Strong, Adequate, Weak, or Missing. Must match score. |
confidence | enum | yes | high, medium, or low. How much evidence supported the assessment. |
assessment | string | yes | 2–3 sentence assessment (same content as the prose review). |
evidence | array | yes | At least one evidence entry per dimension. |
recommendation | string/null | yes | Null if label is Strong. Required otherwise. |
Bad-strategy pattern objects
| Field | Type | Required | Description |
|---|---|---|---|
pattern | enum | yes | One of: fluff, failure_to_face_challenge, goals_as_strategy, bad_objectives, strategy_by_analogy. |
detected | boolean | yes | Whether the pattern was found. |
severity | enum/null | yes | null if detected: false. One of: critical, serious, moderate when detected. |
evidence | array | yes | Evidence entries. Empty array if detected: false. |
resolution | string/null | yes | Suggested fix. null if detected: false. |
Interview lens audit objects
| Field | Type | Required | Description |
|---|---|---|---|
lens | string | yes | Lens identifier (e.g., landscape_mapping, choice_cascade, value_innovation). |
assessment | enum | yes | One of: reflected, partially_reflected, missing. |
survived | array[string] | yes | Specific findings that made it into the draft. May be empty. |
dropped | array[string] | yes | Specific findings lost between notes and draft. May be empty. |
impact_on_kernel | string | yes | How the lens findings affected the kernel elements. |
evidence | array | yes | Evidence entries with provenance. |
Strengths objects
Each entry in the optional top-level strengths array captures one specific strength the author should protect. Mirrors the prose "What Works" section. When emitted, the array must contain 2–4 entries.
| Field | Type | Required | Description |
|---|---|---|---|
title | string | yes | Short, specific title for the strength. |
description | string | yes | 1–2 sentences on why it works and what it anchors. |
evidence | array | yes | At least one evidence entry pointing to the supporting passage. |
Notes cross-reference object
The optional top-level notes_cross_reference object captures structured signals from comparing strategy-notes.md (or interview durable state) against the draft. Only emit when notes were available to the reviewer. Mirrors the prose "Notes Cross-Reference" section.
| Field | Type | Required | Description |
|---|---|---|---|
patterns_that_crept_back | array | no | Bad-strategy patterns caught during interview that reappeared in the draft, often in softer language. May be empty. |
thinking_sharpened_then_softened | array | no | Topics where the notes show a sharper version of the diagnosis or policy than the draft contains. May be empty. |
patterns_that_crept_back entries:
| Field | Type | Required | Description |
|---|---|---|---|
pattern | enum | yes | One of the five bad-strategy patterns. Same vocabulary as bad_strategy_patterns. |
description | string | yes | What the pattern looks like in the softer draft form. |
notes_quote | string | yes | The original sharper language from the notes. |
draft_quote | string | yes | The softened version that appears in the draft. |
thinking_sharpened_then_softened entries:
| Field | Type | Required | Description |
|---|---|---|---|
topic | string | yes | What was softened — typically diagnosis, guiding_policy, or a named action. |
notes_version | string | yes | The sharper formulation from the notes. |
draft_version | string | yes | The softened formulation in the draft. |
what_was_lost | string | yes | The specific edge, claim, or admission that the softening removed. |
Note: notes_cross_reference complements but does not duplicate interview_lens_audit (which captures lens-specific drift) or unresolved_questions (open questions). When the same finding is best expressed as a lens audit entry, prefer interview_lens_audit.
Recommended next steps objects
Each entry in the optional top-level recommended_next_steps array is one priority action. Array order encodes priority — first entry is highest priority. Mirrors the prose "Recommended Next Steps" section. When emitted, the array must contain 2–4 entries.
| Field | Type | Required | Description |
|---|---|---|---|
action | string | yes | Concrete action the author should take. Specific enough to act on without asking "but what?". |
rationale | string | yes | Why this is the priority — what it addresses, what it unblocks. |
Aggregate score
| Field | Type | Description |
|---|---|---|
weighted_total | number | Σ(dimension_score × weight / 100). Range: 1.0–4.0. |
weights | object | Percentage weights per dimension. Must sum to 100. |
label | enum | Overall label derived from weighted_total: Strong (≥3.5), Adequate (≥2.5), Weak (≥1.5), Missing (<1.5). |
confidence | enum | high, medium, or low. Reflects the lowest confidence among the three highest-weighted dimensions. |
Default weights:
| Dimension | Weight | Rationale |
|---|---|---|
| Diagnosis quality | 20 | Foundation — everything downstream depends on it |
| Guiding policy strength | 20 | The core strategic choice |
| Action coherence | 15 | Execution readiness |
| Kernel chain integrity | 15 | Overall logical coherence |
| Bad-strategy patterns | 10 | Anti-pattern detection |
| Assumption exposure | 10 | Risk awareness |
| Specificity / falsifiability | 10 | Testability and accountability |
Reward signal
| Field | Type | Description |
|---|---|---|
pass | boolean | true if weighted_total ≥ 2.5 AND blocking_findings is empty. |
score | number | Normalized 0.0–1.0: (weighted_total - 1.0) / 3.0. |
rationale | string | One-sentence explanation of pass/fail. Must cite blocking findings if pass is false. |
Pass/fail logic:
- A strategy passes if its aggregate is at least Adequate AND no findings are blocking.
- A critical finding with severity "critical" that affects a load-bearing element (diagnosis, guiding policy, or the load-bearing action) is blocking by default.
- The reviewer may mark additional findings as blocking based on judgment.
Failure path pattern vocabulary
Use these values for the pattern field in failure_paths. These cover the most common failure modes. If a scenario doesn't fit any of these, use a descriptive snake_case value — but prefer these standard patterns for pipeline consistency:
| Pattern | Description |
|---|---|
capability_gap | Load-bearing action requires capability the org doesn't have |
slow_drift | Guiding policy exclusions erode through accumulated exceptions |
unmodeled_response | Strategy assumes static competitive environment |
political_veto | Strategy requires stopping something a powerful stakeholder owns |
assumption_cascade | One failed assumption propagates through tightly coupled actions |
execution_bottleneck | Multiple actions depend on same scarce resource |
measurement_void | No leading indicators; problems invisible until too late |
---
Validation rules
Before emitting the JSON, verify:
1. Parseable: Output is valid JSON. No trailing commas, no comments, no markdown fencing in the file itself. 2. Required fields present: All fields marked required in the field reference exist. 3. Score-label consistency: Every score matches its label per the mapping table. A score: 3 with label: "Weak" is invalid. 4. All seven dimensions present: dimension_ratings has exactly seven entries, dimensions 1–7. 5. All five patterns present: bad_strategy_patterns has exactly five entries, one per pattern. 6. Evidence non-empty: Every dimension rating has at least one evidence entry. Every critical finding has at least one evidence entry. 7. Weighted total correct: aggregate_score.weighted_total equals Σ(score × weight / 100) within ±0.01. 8. Aggregate label correct: Label matches the weighted_total per the threshold table. 9. Reward signal consistent: pass is true only if weighted_total ≥ 2.5 AND blocking_findings is empty. 10. Blocking findings reference real findings: Every entry in blocking_findings matches a title in critical_findings. 11. Provenance valid: Every provenance value is one of the five allowed enum values. 12. Recommendations present when needed: recommendation is non-null for any dimension with label other than Strong. 13. Confidence values valid: Every confidence value is one of: high, medium, low. 14. Bad-strategy pattern names valid: The five bad_strategy_patterns entries use exactly these pattern values: fluff, failure_to_face_challenge, goals_as_strategy, bad_objectives, strategy_by_analogy. 15. Failure paths minimum count: failure_paths has at least 2 entries. 16. Critical findings count: critical_findings has 2–4 entries. 17. Strengths count when present: If strengths is emitted, it has 2–4 entries (matching the prose template). Each entry has at least one evidence entry. 18. Notes cross-reference precondition: notes_cross_reference is emitted only when strategy-notes.md or interview durable state was available to the reviewer AND at least one sub-array (patterns_that_crept_back or thinking_sharpened_then_softened) is non-empty. Omit the field rather than emitting an empty object or an object with two empty arrays. 19. Patterns-that-crept-back vocabulary: Every notes_cross_reference.patterns_that_crept_back[].pattern value uses the same five-pattern vocabulary as bad_strategy_patterns. 20. Recommended next steps order: recommended_next_steps array order encodes priority — first entry is highest priority. Each entry has both action and rationale. 21. Recommended next steps count when present: If recommended_next_steps is emitted, it has 2–4 entries.
---
Relationship to prose review
The JSON artifact is a structured projection of the prose review, not a replacement. The two must be consistent:
- Every dimension label in the JSON matches the rating in
strategy-review.md. - Every critical finding in the JSON appears in the prose Critical Findings section.
- The aggregate label in the JSON matches the overall assessment tone in the Review Summary.
- Evidence quotes in the JSON are verifiable against the same document passages cited in the prose.
Section-to-field mapping (when the optional fields are emitted):
Prose section (review-template.md) | JSON field |
|---|---|
| What Works | strengths |
| Critical Findings | critical_findings |
| Dimension Ratings | dimension_ratings |
| Failure Path Analysis | failure_paths |
| Assumption Risk Map | assumption_risk_map |
| Lens Findings (review lenses) | review_lenses_applied |
| Lens Findings (interview lens audit subsection) | interview_lens_audit |
| Notes Cross-Reference — Patterns that crept back | notes_cross_reference.patterns_that_crept_back |
| Notes Cross-Reference — Thinking sharpened then softened | notes_cross_reference.thinking_sharpened_then_softened |
| Notes Cross-Reference — Open questions still open | unresolved_questions |
| Notes Cross-Reference — Lens findings not reflected | interview_lens_audit (with dropped populated) |
| Recommended Next Steps | recommended_next_steps |
When both are produced, write the prose first, then derive the JSON from it. This ensures the reviewer's judgment flows from careful prose analysis, not from filling in a schema.
---
What downstream consumers can expect
External frameworks (evaluation pipelines, RL loops, comparison dashboards) can rely on:
- Stable field names within a
schema_version. Breaking changes increment the version. - Complete dimension coverage: All seven dimensions, all five bad-strategy patterns, always present.
- Evidence provenance: Every score traces to tagged evidence. Consumers can filter by provenance to distinguish document facts from reviewer inferences.
- Deterministic pass/fail: The
reward_signal.passfield follows a documented formula — no hidden judgment. - Blocking findings as a gate: A pipeline can use
blocking_findings.length === 0as a quality gate.
Pressure-test scenarios
Expected behaviors for the strategy-review skill. Use these to validate that the skill handles common review entry points correctly.
| Scenario | Expected behavior |
|---|---|
| User provides a standalone goals document (all aspirations, no diagnosis) | Name it: "This reads more like a goals document than a strategy." Rate Diagnosis as Missing, Bad-Strategy Patterns as Weak (goals masquerading as strategy). Offer to help identify the missing strategic elements. |
User provides strategy-draft.md only, notes exist but weren't offered | Ask for strategy-notes.md — the notes dramatically enrich the review. Proceed without if user declines. |
| Draft + notes mismatch: notes contain sharper thinking than draft | Flag in Notes Cross-Reference: "The interview produced a sharper diagnosis than the draft contains." Quote both versions. |
| Polished board deck with weak diagnosis | Calibrate to near-final maturity — higher bar, flag everything that undermines credibility. Focus on the diagnosis gap specifically because a polished presentation with a vague diagnosis is the most dangerous kind of bad strategy. |
| Comparative review: two competing proposals | Review each independently first using the seven dimensions, then compare. Don't declare a winner unless structural quality clearly differs. |
Mid-interview provisional draft ([PROVISIONAL] tag) | Produce a "mid-interview check" with 3-5 observations, not a full review. Focus on strongest and weakest emerging kernel elements. Suggest questions to pressure-test before finalizing. |
| Strategy with sourced market claims | Check which claims are sourced vs. asserted. Flag conventional wisdom presented as fact. Note claims you can't verify: "I can't verify this, but if wrong, the strategy changes." If durable state is active, tag claims in source-evidence.md. |
| Long multi-appendix document (>3,000 words) | Activate durable review state. Extract evidence to source-evidence.md during reading. Compose final review from artifacts, not conversation memory. |
| User says "poke holes in this" or "quick take" | Deliver findings inline in chat — lead with Critical Findings, top failure path, recommendations. Offer to write full strategy-review.md afterward. Do not write the file without confirming. |
| Strategy already in execution (post-hoc review) | Focus on assumption validation — which assumptions can now be checked against reality? Flag any already falsified by events. |
| Strategy uses OKRs/V2MOM/SWOT instead of kernel language | Evaluate the thinking, not the format. Map concepts to kernel. Flag genuine structural gaps ("this V2MOM has methods but no diagnosis") rather than forcing vocabulary. |
| Notes contain Wardley findings (evolution stages, value-chain dependencies, inertia insight); draft drops them | Flag in interview lens audit: "Landscape mapping — Partially reflected" or "Missing." Identify the specific findings that were dropped — e.g., notes say "our data pipeline is custom-built but three vendors now offer this as a product" but the draft says "we face competitive pressure in infrastructure." Quote both versions. Record under Dim 1 (diagnosis quality) because the landscape insight was meant to sharpen the diagnosis. Do NOT substitute Five Forces questions to validate the Wardley findings — Five Forces tests competitive pressure, not value-chain mapping fidelity. |
| Notes contain cascade capability gaps; draft treats them as action items | Flag in interview lens audit: "Strategic choice cascade — Partially reflected." The cascade identified a capability the org doesn't have as a risk; the draft treats it as a task ("hire mobile team by Q3"). Record under Dim 6 (assumption exposure) — the gap between "we need this capability" and "we have this capability" is a load-bearing assumption. |
Judge artifact mode scenarios
| Scenario | Expected behavior |
|---|---|
| Normal review — user says "review this strategy" with no judge/JSON language | Produce only strategy-review.md (or chat-only feedback). Do NOT emit strategy-review.json. No mention of JSON artifacts unless the user asks. |
| User says "produce a machine-readable review" or "LLM judge output" | Activate judge artifact mode. Produce strategy-review.md first, then strategy-review.json. JSON must parse as valid JSON, contain all seven dimension ratings and all five bad-strategy pattern entries, and every score must match its label per the mapping (Strong=4, Adequate=3, Weak=2, Missing=1). |
| User says "JSON only" or "just the evaluation artifact" | Produce only strategy-review.json, skip the prose file. All validation rules still apply. |
| Judge mode review of a strategy with missing diagnosis | JSON dimension_ratings[0] (diagnosis_quality) must have score: 1, label: "Missing". The critical_findings array must include a finding about the missing diagnosis. That finding's title must appear in blocking_findings. The reward_signal.pass must be false because blocking findings are non-empty. The rationale must cite the missing diagnosis. |
| Judge mode review where interview notes contain Wardley lens findings that were dropped from the draft | JSON interview_lens_audit must include an entry with lens: "landscape_mapping" and assessment: "partially_reflected" or "missing". The dropped array must name the specific findings lost between notes and draft. The evidence array must contain at least one notes_cross_ref entry (the original finding) and one doc_quote entry (what replaced it or the absence). The prose review's Lens Findings section must contain the same audit finding. |
| Judge mode — weighted total and reward signal arithmetic | aggregate_score.weighted_total must equal Σ(score × weight / 100) within ±0.01. reward_signal.score must equal (weighted_total - 1.0) / 3.0 within ±0.01. reward_signal.pass must be true only when weighted_total ≥ 2.5 AND blocking_findings is empty. |
| Judge mode — prose and JSON consistency | Every dimension label in JSON must match the rating in strategy-review.md. Every critical finding title in JSON must appear in the prose Critical Findings section. If the prose says "Diagnosis Quality — Weak" but the JSON has score: 3, label: "Adequate", that is a validation failure. |
| Judge mode — adequate aggregate with blocking findings | Strategy scores Adequate or above on aggregate (weighted_total ≥ 2.5) but has a critical finding in blocking_findings. The reward_signal.pass must be false despite the adequate aggregate. The rationale must cite the blocking finding, not the score. This validates that blocking findings gate independently of the aggregate — the AND condition in the pass/fail logic requires BOTH adequate score AND empty blocking findings. |
| Judge mode — happy path (all dimensions Strong) | Strategy is structurally sound on every dimension. All seven dimension_ratings have score: 4, label: "Strong". Every dimension's recommendation is null. No bad_strategy_patterns entries are detected (detected: false for all five). blocking_findings is empty. aggregate_score.weighted_total is 4.0 with label: "Strong". reward_signal.pass is true and reward_signal.score is 1.0. critical_findings still has 2–4 entries — even a strong strategy has soft spots — but each is severity moderate at most and none appear in blocking_findings. If strengths is emitted, it has 2–4 entries. The prose Review Summary must not bury or invent flaws to balance the assessment; if the strategy is genuinely strong, say so. |
| Judge mode — durable review state active with judge artifact mode | User invoked judge artifact mode on a long multi-appendix document for which durable review state was activated (per the "Long multi-appendix document" scenario in the table above). The skill must produce strategy-review.md and strategy-review.json in the user's working directory AND also write strategy-review.json to the durable review state directory. Both copies must be byte-identical. If strategy-notes.md was present, the JSON must include the notes_cross_reference object. If interview lenses appear in the notes, the JSON must include interview_lens_audit with evidence entries that include at least one notes_cross_ref provenance tag pointing into the durable state files. The review_metadata.review_id should be emitted so pipeline consumers can deduplicate the dual-written artifact. Validation: read both copies and confirm the dual-write happened and that they match. |
| Judge mode — partial-write failure during dual-write | One of the two write locations succeeds while the other fails (disk full, permission denied, path not found). The skill must warn about the failed write location but still complete the review — the successful copy is usable. The warning must name which location failed and include the error. Do not retry or roll back the successful write. |
Review Dimensions
Seven dimensions for evaluating strategy documents. Each dimension has criteria for each rating level, common failure patterns, and the kinds of pressure-test questions to apply.
The overall goal: find where the strategy could break — the gaps, unexamined risks, and failure paths that the author hasn't accounted for. A strategy that survives this review has a meaningfully better chance of surviving contact with reality.
---
1. Diagnosis Quality
Core question: Does the diagnosis name the actual challenge with enough specificity that you could imagine being wrong about it?
A diagnosis isn't a description of the situation — it's a judgment about what matters most. It picks from a complex reality and says "this is the thing." That act of picking is what makes it useful and what makes it testable.
Rating criteria
Strong: Names a specific challenge. Uses the company's particular situation, not generic industry observations. You could argue against this diagnosis — it takes a position. Often uses analogy or identifies a structural pattern ("this is a classic disintermediation problem," "we're in the position Kodak was in when digital hit 5% market share").
Adequate: Identifies a real challenge but stays somewhat generic. Could apply to several companies in the same industry without modification. Correct but not sharp enough to generate a distinctive guiding policy.
Weak: Describes the situation without identifying what matters most. Lists multiple challenges without prioritizing. Uses language like "the market is changing" or "we face increasing competition" without specifying how or why it matters for this company specifically.
Missing: No diagnosis at all. The document jumps from context to goals or actions. The most common version: a SWOT analysis with no "therefore."
Pressure-test questions
- Could a competitor read this diagnosis and say "yes, that's exactly our challenge too"? If so, it's not specific enough.
- Does the diagnosis identify why now? What changed that makes this challenge urgent or this opportunity available?
- Is there an unstated assumption about the external environment that, if wrong, would invalidate the diagnosis entirely?
- Does the diagnosis identify the root challenge, or a symptom? ("We're losing market share" is a symptom. "Our product architecture prevents us from serving the mobile-first workflow that 68% of inspections now use" is a root cause.)
- What would the diagnosis look like if the author's biggest fear were true? What if their most optimistic assumption were false?
Common failure patterns
- The comfortable diagnosis: Identifies a challenge that the organization already knows how to solve, avoiding the harder truth that requires a different approach.
- The diagnosis-by-data-dump: Presents extensive market research and trend analysis but never synthesizes it into a judgment. The reader is left to figure out what it all means.
- The inherited diagnosis: Repeats the diagnosis from last year's strategy without checking whether the underlying situation has changed.
- The consensus diagnosis: So carefully worded to avoid offending any stakeholder that it says nothing specific. Usually produced by committee.
Lens connection — Competitive Pressure Audit: When the diagnosis addresses competitive dynamics, apply Five Forces thinking from references/review-lenses.md. Check whether the diagnosis accounts for all five forces (rivalry, new entrants, substitutes, supplier power, buyer power) or focuses narrowly on direct rivalry while ignoring structural pressures that could reshape the competitive environment. A diagnosis that says "we're losing to Competitor X" but doesn't consider substitution risk or eroding entry barriers has only seen part of the picture.
---
2. Guiding Policy Strength
Core question: Does the guiding policy make a real choice that rules things out — including things a reasonable person might choose?
A guiding policy that everyone agrees with hasn't decided anything. The value is in what it excludes. If a competitor could adopt the same policy verbatim without contradiction, it's not doing strategic work.
Rating criteria
Strong: Makes a clear directional choice. Names what the organization will not do. The rejected alternatives are plausible — someone could reasonably argue for them. Exploits a specific asymmetry or pivot point. Short enough to remember.
Adequate: Makes a choice but the exclusions are implicit rather than stated. You can infer what's been ruled out but it hasn't been made explicit, which means different readers may interpret the scope differently.
Weak: Describes a desirable direction without excluding anything. Could be adopted by any company in the industry. Language like "focus on the customer," "drive innovation," "be the best" — these sound like policies but aren't, because they don't constrain decisions.
Missing: No guiding policy. The document has goals and actions but nothing connecting them — no principle that explains why these actions and not others.
Pressure-test questions
- What does this policy prevent the organization from doing? Can the author name three specific initiatives or investments they would decline because of this policy?
- Is there a real alternative to this policy that a thoughtful person might advocate? If not, it's not a choice — it's a platitude.
- Does the policy address the diagnosis? Trace the line: the diagnosis says X is the challenge, the policy says we'll address it by Y. Does Y actually respond to X, or is it a non sequitur?
- Who in the organization would be uncomfortable with this policy? If nobody, it hasn't decided anything.
- What does this policy look like in a bad quarter? Does it still hold when resources are tight, or would the organization quietly abandon it?
Common failure patterns
- The umbrella policy: Broad enough to accommodate any initiative anyone wants to pursue. "Invest in growth" covers everything from R&D to acquisitions to sales hiring — it constrains nothing.
- The stealth goal: "Our policy is to achieve market leadership in segment X" — that's a goal wearing a policy costume. The policy would be the approach to achieving market leadership: through price, through service depth, through platform lock-in, etc.
- The both/and policy: "We will be the cost leader AND the innovation leader AND the customer service leader." Strategies that refuse to choose between contradictory positions are not strategies.
- The re-emerging rejected alternative: The guiding policy explicitly rules something out, but one of the coherent actions quietly reintroduces it. Check for this specifically.
---
3. Action Coherence
Core question: Do the actions reinforce each other, and would they actually carry out the guiding policy?
Coherence is the tell. An incoherent action set — actions that pull in different directions, duplicate effort, or silently contradict each other — is the signature of a strategy that was assembled by committee rather than designed as a system.
Rating criteria
Strong: Actions clearly reinforce each other — doing A makes B easier, B makes C cheaper, C protects A. The set has focus (3-6 actions, not 15). Each action has an owner, a timeframe, and a clear connection to the guiding policy. The author can trace the chain: if you remove any one action, the system degrades.
Adequate: Actions are consistent with the guiding policy and don't contradict each other, but the reinforcement is weak or unstated. They're a reasonable to-do list for someone pursuing this policy, but they haven't been designed as a system.
Weak: Actions include items that contradict or compete with each other for resources. Or the action list is too long (8+) and lacks prioritization — it's a project portfolio, not a coherent set. Or actions are so vague ("invest in technology") that coherence can't be assessed.
Missing: No actions, or actions that are disconnected from both the diagnosis and the guiding policy. The document ends with the policy and leaves execution to "the teams."
Pressure-test questions
- For each pair of actions: does A make B easier or harder? Are there hidden resource conflicts?
- Which action is load-bearing? If it fails, does the rest of the strategy still work, or does it collapse? This identifies single points of failure.
- Are there actions that would happen regardless of this strategy — business-as-usual relabeled as strategic? Remove those and see what's left.
- Do the actions require capabilities the organization doesn't have? If so, is building those capabilities itself one of the actions?
- What's the sequence? Some actions must precede others. If the document treats them as parallel but they're actually sequential, the timeline is fiction.
- What happens if the most expensive or most difficult action is delayed by six months? Does the strategy degrade gracefully or break?
Common failure patterns
- The laundry list: Every team's top priority became a "strategic action." No prioritization, no sequencing, no resource allocation.
- The invisible dependency: Action 3 depends on action 1 being complete, but this isn't stated. The actions look parallel but are actually sequential, and the timeline doesn't account for it.
- The orphan action: One action doesn't connect to any other. It was probably added to satisfy a stakeholder and wasn't part of the original strategic logic.
- The missing action: The guiding policy implies something that doesn't appear in the action list. For example, the policy says "focus on segment X" but no action addresses exiting segment Y.
Lens connection — Execution Alignment: When the strategy requires organizational change, apply 7S thinking from references/review-lenses.md. Actions that demand new capabilities, cross-functional coordination, or cultural shifts carry invisible requirements the strategy may not account for — check whether the actions address alignment across structure, systems, shared values, style, staff, and skills, or assume the organization will simply follow the new direction.
Lens connection — Strategy Deployment: When the strategy involves multiple organizational levels, apply Hoshin Kanri thinking from references/review-lenses.md. Check whether actions are stated at the right altitude and whether translation from executive intent to team-level execution has been accounted for. If the strategy adds actions without subtracting existing work, execution teams will make their own prioritization choices — and those choices may silently undermine the strategy.
---
4. Kernel Chain Integrity
Core question: Does the logical chain — "because [diagnosis], we will [guiding policy], which means we will [actions]" — actually hold?
This is the meta-dimension. Dimensions 1-3 evaluate each kernel element individually; this one evaluates whether they fit together. A strategy can have a good diagnosis, a good policy, and good actions that don't logically connect.
Rating criteria
Strong: The "therefore" test passes cleanly. Reading "[diagnosis]. Therefore, [policy]. Which means [actions]." sounds like a logical argument, not a slide transition. Each link is tight — the policy responds to the specific diagnosis (not a different challenge), and the actions carry out the specific policy (not a generic version of it).
Adequate: The chain is traceable but requires some inference. The connections are present but not tight — you can see how they relate, but the policy could serve a somewhat different diagnosis, or the actions could serve a somewhat different policy.
Weak: Significant gaps in the chain. The diagnosis and policy address different topics. Or the actions seem disconnected from the policy — they might be good actions for a different strategy.
Missing: The document has components that don't form a chain at all — either because elements are missing, or because they were written independently and assembled without checking fit.
Pressure-test questions
- Read the strategy as a single paragraph: "[Diagnosis]. Therefore, [policy]. Which means [actions]." Does it sound coherent, or do the transitions feel forced?
- If the diagnosis were different — say the opposite were true — would the guiding policy still make sense? If yes, the policy isn't responding to the diagnosis.
- If the guiding policy ruled something out but an action quietly reintroduces it, the chain is broken. Check the exclusion list against each action.
- Could someone read just the actions and correctly guess the diagnosis? If the actions don't contain the DNA of the diagnosis, the chain has broken by the time it reaches execution.
---
5. Bad-Strategy Pattern Detection
Core question: Are any of the four hallmarks of bad strategy (plus the additional anti-pattern) present in the document?
Bad strategy isn't the absence of strategy — it's the presence of specific patterns that look like strategy but aren't. These patterns are dangerous because they create confidence without substance. A document that contains them will feel "strategic" to casual readers while providing no actual guidance for decisions.
The four Rumelt hallmarks (plus one additional anti-pattern)
1. Fluff — Buzzword-heavy language that sounds sophisticated but says nothing.
- Signal phrases: synergy, leverage, ecosystem, platform play, holistic, transformational, best-in-class, next-generation, customer-centric.
- The test: replace the buzzword with its plain-language definition. If the sentence now says nothing, it was fluff.
- Note: a single buzzword in an otherwise concrete document isn't a finding. The pattern is when fluff replaces specificity rather than supplementing it.
2. Failure to face the challenge — No clear statement of what the actual problem is. The document is all aspiration and no obstacle.
- The test: can you state the challenge in one sentence? If you can only state the desired outcome, the challenge has been avoided.
3. Mistaking goals for strategy — Revenue targets, market share goals, or outcome metrics presented as if they were strategy. "Grow 30% YoY" is a wish, not a plan.
- The test: does the document explain how and why that how? If the "strategy" section could be replaced by a spreadsheet target, it's goals.
4. Bad strategic objectives — Either a laundry list with no prioritization (the "dog's dinner") or blue-sky objectives that assume away the hard part ("we will eliminate tech debt").
- The test: if every objective is "high priority," none of them are. If the objective restates the problem as if naming it solved it, it's blue-sky.
5. Strategy by analogy (additional anti-pattern, not a Rumelt hallmark) — "Netflix did X, so we should too" without examining whether the conditions that made it work for Netflix exist here.
- The test: can the author name three conditions that made the analogy work for the original company, and verify they hold here?
Rating criteria
Strong: No hallmarks detected, or minor traces that don't undermine the strategy's substance.
Adequate: One or two hallmarks present in mild form — some fluff in the executive summary, a section that reads more like goals than strategy. The core thinking is sound but the document could be tightened.
Weak: Multiple hallmarks present. The document has significant sections that sound strategic but lack substance. The reader would struggle to make a concrete decision based on this strategy.
Missing (meaning: the entire document is bad strategy): The document is primarily goals, fluff, or laundry lists with no underlying strategic logic. This is a "start over" finding.
---
6. Assumption Exposure
Core question: Has the strategy identified its load-bearing assumptions, and are those assumptions testable?
Every strategy is a bet. The question isn't whether it rests on assumptions — all strategies do — but whether the author knows which assumptions are load-bearing and has a plan for what happens if they're wrong.
Rating criteria
Strong: Load-bearing assumptions are explicitly stated, distinguished from background assumptions. Each one is testable — the author could describe what evidence would confirm or refute it. The strategy acknowledges what changes if the biggest assumption is wrong. Risk exposure is proportional to the stakes.
Adequate: Some assumptions are stated but the list is incomplete. The most obvious ones are there; the subtle or uncomfortable ones are missing. Limited discussion of what happens if assumptions fail.
Weak: Assumptions are either absent or buried in optimistic language ("as the market continues to grow..."). The strategy reads as if the environment is certain and the plan will execute as designed.
Missing: No assumptions stated. The strategy presents its view of the world as fact. This is especially dangerous when the strategy depends on market trends, competitor behavior, or technology adoption curves.
Pressure-test questions
- What are the three assumptions that, if wrong, would break this strategy entirely? Are they stated in the document?
- Does the strategy assume competitor behavior (e.g., "competitors will be slow to respond")? That's almost always optimistic.
- Does the strategy assume capability that doesn't yet exist (e.g., "once we build the new platform")? What happens if it takes twice as long?
- Does the strategy assume market conditions will continue (e.g., "the market is growing at 15% annually")? What if growth stalls?
- Are there second-order effects that haven't been considered? If action A succeeds, what does the competitor do in response? What does the customer do?
- Does the strategy assume organizational willingness? "We will stop doing X" — will the team that owns X agree? Is there political resistance the strategy pretends doesn't exist?
- What's the failure path that doesn't involve a dramatic crisis — the slow drift scenario where the strategy dies quietly because nobody noticed it wasn't working?
Common failure patterns
- Optimism as strategy: The document's assumptions about the future consistently break in the org's favor. Markets grow, competitors stumble, technology works, talent arrives. No downside scenarios.
- The missing "what if": No discussion of what changes if a key assumption fails. A strategy without contingency exposure is a strategy without humility.
- Invisible assumptions: The most dangerous kind. The strategy implicitly assumes things that nobody has stated — for example, that the team has the skills to execute a new approach, or that the board will fund the investment, or that customers will switch behavior.
- The political assumption: The strategy assumes organizational alignment that doesn't exist. "We will consolidate the three product lines" — has anyone told the three product line owners?
Lens connection — Execution Alignment: When the strategy requires organizational change, the most dangerous assumptions are often about the organization itself. Apply 7S thinking from references/review-lenses.md to surface unstated assumptions about whether current structure, systems, culture, and leadership style will support the new direction. The "slow drift" and "political veto" failure paths are frequently 7S misalignment that nobody named as an assumption.
---
7. Specificity and Falsifiability
Core question: Could you tell in 12 months whether this strategy worked, partially worked, or failed? If so, how?
A strategy that can't be evaluated can't be improved. Specificity isn't about adding metrics for the sake of metrics — it's about making the strategy's claims testable. If the diagnosis is vague, you can't tell whether the guiding policy addressed it. If the success criteria are vague, you can't tell whether the actions worked.
Rating criteria
Strong: Each kernel element is specific enough to be falsified. The diagnosis names something concrete that can be verified. The guiding policy's exclusions are clear enough that you'd notice a violation. The actions have enough detail to tell whether they happened. Success indicators are observable, not aspirational.
Adequate: Mostly specific but with pockets of vagueness. The diagnosis is concrete but the success indicators are fuzzy. Or the actions are clear but the guiding policy is broad enough to accommodate almost any action set.
Weak: Significant vagueness throughout. You could read the strategy a year from now and have no way to assess whether it was followed, whether the diagnosis was correct, or whether the actions produced the intended effect.
Missing: The strategy is entirely unfalsifiable — so vague that no outcome could ever contradict it. "We will drive growth through innovation" cannot fail because it never committed to anything specific.
Pressure-test questions
- Imagine you're reviewing this strategy in 12 months. What evidence would tell you it succeeded? What evidence would tell you it failed? If you can't describe the failure case, the strategy isn't specific enough.
- Are the success indicators leading indicators (detect problems early) or lagging indicators (confirm what already happened)? A strategy with only lagging indicators can't course-correct.
- Could a new employee read this document and understand what they should and shouldn't work on? If not, it's not specific enough to guide decisions.
- Take each action: what does "done" look like? Is there a clear difference between "we did this" and "we said we would do this but didn't"?
Lens connection — Objective Translation: When success criteria are one-dimensional or vague, apply Balanced Scorecard thinking from references/review-lenses.md. Check whether the strategy has leading indicators from customer and internal-process perspectives that would signal problems before financial metrics confirm them. A strategy measurable only in revenue and market share has a 6-12 month blind spot between cause and detection.
Review Lenses
Four complementary frameworks that sharpen the review when the strategy document warrants them. Lenses load conditionally — most reviews use one or two; some use none. The reviewer weaves lens-specific questions into the existing seven dimensions rather than adding new sections.
The lenses do not ask the author to adopt a new framework. They give the reviewer sharper questions by applying structured thinking that the core dimensions alone might miss — particularly around execution feasibility and competitive environment accuracy.
---
1. McKinsey 7S — Execution Alignment
When to load
When the strategy requires organizational change — new structure, new processes, cultural shift, or cross-functional coordination. Load this lens when you see:
- Actions that imply structural reorganization (new teams, new reporting lines, consolidated functions)
- A guiding policy that requires different skills or capabilities than the organization currently has
- Cross-functional dependencies that the actions treat as straightforward
- References to "culture change," "mindset shift," or "new ways of working"
What to look for
The seven elements: strategy, structure, systems, shared values, style, staff, skills. The most common gap: the strategy changes the "strategy" element but assumes the other six will follow.
Check alignment between what the strategy demands and what the organization currently is:
- Structure: Does the current org structure support the new direction, or does it create friction? If teams need to collaborate differently, has the strategy addressed how?
- Systems: Do existing processes, tools, and decision-making systems support the change, or will they produce antibodies?
- Shared values: Does the strategy conflict with deeply held organizational beliefs? A strategy requiring "move fast" in an organization that values "measure twice" will face silent resistance.
- Style: Does leadership behavior reinforce the strategy? If the strategy says "empower teams" but leadership is command-and-control, the strategy loses.
- Staff: Does the organization have the people this requires? Not just headcount — the specific mix of experience, judgment, and domain knowledge.
- Skills: Does the organization have the capabilities? Not individual skills — organizational capabilities like "shipping mobile products" or "managing enterprise sales cycles."
Pressure-test questions
1. Which of the six non-strategy elements (structure, systems, shared values, style, staff, skills) must change for this strategy to work? Does the document acknowledge any of them? 2. Where will current systems and processes produce friction against the new direction? Is that friction addressed or ignored? 3. If the strategy requires cross-functional coordination that doesn't exist today, what's the plan for creating it? "Better communication" is not a plan. 4. Does leadership style match what the strategy demands? A strategy requiring decentralized decisions won't work under centralized leadership. 5. What's the gap between the skills the organization has and the skills the strategy needs? Is closing that gap part of the action plan, or is it assumed away?
Common gaps this reveals
- The structural orphan: A key action requires coordination between teams that don't currently work together, and no one owns making that coordination happen.
- The systems antibody: Existing approval processes, budgeting cycles, or performance metrics will actively undermine the new strategy because nobody thought to change them.
- The cultural contradiction: The strategy requires behaviors that conflict with "how we actually do things here." The document won't mention this; the reviewer should.
- The skills gap dressed as a timeline: "Hire a mobile team by Q2" treats a capability-building challenge as a procurement problem.
Dimension connections: Primarily strengthens Dimension 3 (Action Coherence) and Dimension 6 (Assumption Exposure). The "slow drift" and "political veto" failure paths are often 7S misalignment in disguise.
---
2. Balanced Scorecard — Objective Translation
When to load
When the strategy's success criteria are one-dimensional or vague. Load this lens when you see:
- Success measured only in financial terms (revenue, margin, market share)
- "What success looks like" is aspirational rather than observable
- No leading indicators — only outcomes that take 6-12 months to appear
- Actions without clear measures of progress
What to look for
Check whether the strategy translates into objectives across four perspectives — not as a formal scorecard, but as a completeness check on how success is defined:
- Financial: Revenue, margin, cost targets. Almost always present. The question is whether these are the only measures.
- Customer: What changes for the customer? How would you know the strategy is working from the customer's perspective before financial results arrive? Acquisition cost, retention, NPS, usage patterns, time-to-value.
- Internal Process: What processes must improve or be created? How do you measure whether the operational capabilities the strategy requires are actually being built?
- Learning & Growth: Is the organization developing the capabilities, knowledge, and culture the strategy needs? Staff development, knowledge management, innovation pipeline health.
A strategy with only financial success criteria can't course-correct — by the time revenue tells you something's wrong, the underlying causes are 6-12 months old.
Beyond the four perspectives, check for the deeper Balanced Scorecard elements that separate rigorous objective translation from surface-level "balanced KPI coverage":
- Causal linkage: Do the objectives across perspectives form a cause-and-effect chain? Learning & Growth improvements should drive Internal Process improvements, which should drive Customer outcomes, which should drive Financial results. If the objectives are just four independent lists, the scorecard thinking is decorative — it hasn't identified the causal theory of how the strategy creates value.
- Targets: Are there specific, time-bound targets for each objective — not just metrics to watch, but thresholds that distinguish success from failure? "Track NPS" is monitoring; "NPS above 45 by Q3" is a target the org can rally around and course-correct against.
- Strategic initiatives: For each objective gap (where-we-are vs. target), is there a funded initiative to close it? Objectives without initiatives are wishes. Check whether the coherent actions map to specific objective gaps or float independently.
- Review cadence: Does the strategy specify how often progress will be reviewed and by whom? The Balanced Scorecard's operational power comes from regular strategy review meetings where leading indicators trigger mid-course corrections. Without a review rhythm, the scorecard degrades into a dashboard nobody checks.
Pressure-test questions
1. If the strategy is working as intended, what would you see in 90 days that isn't a financial metric? If there's no answer, the strategy has no early warning system. 2. What customer behavior would change first if the strategy is succeeding? Is that behavior being measured? 3. Which internal processes are load-bearing for this strategy? How would you know if they're improving or degrading? 4. What capabilities must the organization build? How would you measure progress on capability-building before end results arrive? 5. If the financial targets are missed in Q3, what non-financial indicators from Q1-Q2 should have flagged the problem? If none exist, the strategy is flying blind. 6. Can you trace a causal chain from a Learning & Growth objective through Internal Process and Customer objectives to a Financial outcome? If the objectives don't link causally, they're four independent wish lists. 7. For each success indicator — is there a specific target with a deadline, or just a metric to "track"? Monitoring without thresholds can't trigger corrective action. 8. For each gap between current state and target — is there a funded initiative to close it, or is the gap expected to close on its own? 9. How often will the strategy's leading indicators be reviewed, by whom, and what authority do they have to adjust course? If there's no review cadence, the scorecard is a one-time artifact, not an operating tool.
Common gaps this reveals
- The financial-only trap: Success defined entirely in outcomes (revenue, market share) with no visibility into the drivers. Problems are discovered 6+ months after they start.
- The activity-as-progress illusion: Actions have timelines but no outcome measures. "Launch mobile app by Q3" tells you whether the action happened, not whether it worked.
- The missing learning loop: No measures for capability development. The strategy requires new skills but has no way to tell whether the organization is acquiring them.
- The customer blind spot: Financial projections assume customer behavior changes, but no customer-facing metrics would confirm whether those changes are happening.
Dimension connections: Primarily strengthens Dimension 7 (Specificity and Falsifiability). Also informs the "measurement void" failure path.
---
3. Porter's Five Forces — Competitive Pressure Audit
When to load
When the diagnosis addresses competitive dynamics. Load this lens when you see:
- Market position or competitive standing as part of the diagnosis
- Pricing pressure, margin erosion, or commoditization discussed
- New market entrants or emerging competitors mentioned
- Technology substitution or category disruption referenced
- Supply chain dependencies or customer concentration acknowledged
What to look for
Check whether the diagnosis has identified the full set of competitive pressures, not just the most visible one:
1. Rivalry among existing competitors: Most strategies address this. The question is whether the analysis goes beyond "we compete with X" to understand structural drivers of rivalry — industry growth rate, differentiation levels, switching costs, exit barriers.
2. Threat of new entrants: What barriers prevent new competitors from entering? Are those barriers eroding? Strategies that assume today's competitor set is fixed often miss that the real threat is someone who doesn't compete today.
3. Threat of substitutes: Not just direct competitors — alternative approaches to the same customer need. The strategy might be winning against Competitor X while a substitute technology makes the entire category irrelevant.
4. Bargaining power of suppliers: How dependent is the organization on specific suppliers, platforms, or partners? A strategy built on a platform you don't control carries supplier-power risk.
5. Bargaining power of buyers: How much leverage do customers have? Strategies that assume pricing power should check whether buyer concentration or switching-cost dynamics support that assumption.
Pressure-test questions
1. The diagnosis addresses rivalry — but has it considered which other forces shape the competitive environment? Which forces are notably absent from the analysis? 2. What prevents a new entrant from pursuing the same guiding policy? If the answer is "nothing," the strategy's advantage window may be shorter than assumed. 3. Is there a substitute technology, business model, or approach that could make the strategy's value proposition irrelevant — not by competing directly, but by redefining the problem? 4. Does the strategy depend on a platform, supplier, or partner whose interests might diverge? What happens if they change terms, compete directly, or disappear? 5. Does the strategy assume customer loyalty or switching costs that may not hold? What happens if buyers gain more leverage through alternatives or consolidation?
Common gaps this reveals
- The rivalry-only diagnosis: Competitive dynamics treated as a two-player game (us vs. main competitor) while ignoring structural forces shaping the entire environment.
- The invisible substitute: Direct competitors correctly identified, but a different category of solution is about to absorb the customer need.
- The platform dependency: Strategy built on a platform (cloud provider, app store, marketplace) treated as infrastructure but actually a supplier with increasing bargaining power.
- The eroding moat: Strategy assumes barriers to entry that are weakening — regulatory protection being removed, proprietary technology being commoditized, talent becoming more available.
Dimension connections: Primarily strengthens Dimension 1 (Diagnosis Quality). For each force the strategy mentions, check the analysis depth. For forces it doesn't mention, ask whether the omission is justified or a blind spot.
---
4. Hoshin Kanri — Strategy Deployment
When to load
When the strategy involves multiple organizational levels — when actions need to cascade from executive intent to team-level execution. Load this lens when you see:
- Actions stated at different altitudes (some executive-level, some operational)
- Multiple teams or departments responsible for different parts of execution
- Initiatives that require translation from strategic intent to operational plans
- A gap between the strategy's ambition and the specificity of its actions
What to look for
Strategy fails at the translation layers. The board-level strategy becomes a VP-level initiative becomes a team-level OKR, and at each translation, meaning is lost:
- Altitude consistency: Are actions stated at the right level? If they're all executive-level ("launch mobile app"), do they account for the team-level translation needed? If they're all operational ("implement offline sync"), is there a clear line back to strategic intent?
- Catchball evidence: Has the strategy been tested bidirectionally — top-down intent meeting bottom-up feasibility? If it reads as purely top-down, the actions may not survive contact with the teams who must execute them.
- Translation gaps: Between each organizational level, ask: could a team lead read this and know what to prioritize this quarter? If not, translation work is needed that the strategy doesn't acknowledge.
- Breakthrough vs. operational: Does the strategy distinguish between breakthrough objectives (change the trajectory) and operational objectives (keep the business running)? Breakthrough objectives need protection from being crowded out by operational demands.
Pressure-test questions
1. If a team lead reads this strategy, can they derive their team's priorities for the next quarter? If not, where does translation break down? 2. Is there evidence that the actions were validated with the people who must execute them — or do they read as mandates from above? 3. Which actions are "breakthrough" (change the trajectory) and which are "operational" (keep the business running)? Are the breakthrough actions protected from being deprioritized when operational demands surge? 4. If the strategy requires coordination between teams, who owns the integration points? If nobody does, those are the points where execution will fragment. 5. What gets deprioritized to make room for these strategic actions? If the strategy adds without subtracting, execution teams will make their own subtraction choices — and those choices may undermine the strategy.
Common gaps this reveals
- The altitude mismatch: Executive-level actions with no operational translation. "Enter the enterprise market" is a strategy-level statement; the teams need to know which segment, through which channel, with which product changes.
- The missing catchball: Actions designed without input from executors. The strategy assumes feasibility that hasn't been validated. This surfaces as "we had no idea this was coming" when strategy is announced.
- The crowded-out breakthrough: Breakthrough objectives that compete with operational demands for the same resources. Without explicit protection, operational urgency always wins — and the strategy quietly dies.
- The orphaned integration: Actions that depend on cross-team coordination, but no one owns the coordination itself. Each team optimizes their piece; the whole fragments.
Dimension connections: Primarily strengthens Dimension 3 (Action Coherence). Also feeds the "execution bottleneck" failure path.
---
Using lenses in the review
Selection
After reading the strategy document and extracting the kernel (Step 2), mentally check each lens trigger. Load the relevant lenses for Step 3. Most reviews activate one or two; activating all four suggests the strategy has broad structural gaps.
Integration
Lens questions weave into the existing seven dimensions — they don't create separate evaluation sections. When a lens reveals a gap, record the finding under the relevant dimension (usually 1, 3, 6, or 7). The lens is the tool that found the gap; the dimension is where the gap lives.
Reporting
If any lenses were applied, include a "Lens Findings" section in the review output (see review-template.md). This section captures what the lenses revealed that the core dimensions alone would likely have missed — particularly systemic issues that span multiple dimensions.
What lenses don't do
- They don't ask the author to adopt a new framework or fill out a grid
- They don't add new dimensions to the review
- They don't replace the core kernel analysis
- They don't require the reviewer to explain the framework to the user — use the thinking, not the vocabulary
Review Output Template
For formal or file-output reviews, produce one file: strategy-review.md. Write it in the user's working directory unless they specify otherwise. If the user asked for quick feedback or a chat-only take, deliver findings inline instead — see the chat-only branch in SKILL.md Step 5. This template applies only when file output is confirmed.
The tone is direct and specific. Every finding must point to evidence in the document. Every recommendation must be concrete enough that the author can act on it without asking "but what specifically should I do?"
---
# Strategy Review: [subject from the strategy document]
_Review of [document name(s)]. Reviewed on [date]._
## Review Summary
[3-4 sentences. What this strategy is trying to do, the overall assessment of its structural integrity, and the single most important thing that needs attention. Don't bury the lede — if the strategy has a critical structural flaw, say it here.]
## What Works
[Specific strengths. Name what the author should protect and build on. This isn't a politeness section — genuinely good elements help anchor what "good" looks like for the rest of the review. 2-4 bullet points, each with a specific citation from the document.]
- **[Strength]**: [Why it works, with reference to the specific passage.]
## Critical Findings
_The 2-4 issues that most undermine the strategy's integrity. These should be addressed before the strategy is shared or acted on. Ordered by severity — most critical first. Lead with these because they are the highest-value part of the review._
### Finding 1: [Title — short, specific]
**Severity:** [Critical / Serious / Moderate]
**What's wrong:** [Specific description of the gap, risk, or failure path.]
**Why it matters:** [What goes wrong in execution if this isn't addressed.]
**Evidence:** [Passage from the document that demonstrates the issue.]
**Recommended fix:** [Concrete action to address it.]
### Finding 2: [Title]
...
## Dimension Ratings
_Supporting detail behind the critical findings. Each dimension evaluates one aspect of the strategy's structural integrity._
### 1. Diagnosis Quality — [Strong / Adequate / Weak / Missing]
**Assessment:** [2-3 sentences on what the diagnosis does well or poorly.]
**Evidence:** [Quote or cite the relevant passage.]
**Recommendation:** [If Adequate or below — specific guidance on what to change. Skip for Strong.]
### 2. Guiding Policy Strength — [Strong / Adequate / Weak / Missing]
**Assessment:** [2-3 sentences.]
**Evidence:** [Quote or cite.]
**Recommendation:** [If needed.]
### 3. Action Coherence — [Strong / Adequate / Weak / Missing]
**Assessment:** [2-3 sentences.]
**Evidence:** [Quote or cite. Name specific action pairs that reinforce or conflict.]
**Recommendation:** [If needed.]
### 4. Kernel Chain Integrity — [Strong / Adequate / Weak / Missing]
**Assessment:** [Read the strategy as "[Diagnosis]. Therefore, [policy]. Which means [actions]." Does it hold?]
**The chain:** [Write it out as a single paragraph. This makes gaps visible.]
**Recommendation:** [If needed. Name the weakest link.]
### 5. Bad-Strategy Patterns — [Strong / Adequate / Weak / Missing]
**Assessment:** [Which patterns, if any, are present.]
**Evidence:** [Quote specific passages for each pattern found. Name the hallmark.]
**Recommendation:** [For each pattern found — how to fix it.]
### 6. Assumption Exposure — [Strong / Adequate / Weak / Missing]
**Assessment:** [Are load-bearing assumptions identified?]
**Unstated assumptions found:** [List assumptions the reviewer identified that the document doesn't state.]
**Recommendation:** [If needed.]
### 7. Specificity and Falsifiability — [Strong / Adequate / Weak / Missing]
**Assessment:** [Could you evaluate this strategy in 12 months?]
**Evidence:** [Point to specific vague or unfalsifiable claims.]
**Recommendation:** [If needed.]
## Failure Path Analysis
_How this strategy could fail — not through dramatic crisis, but through the predictable patterns that kill most strategies. These are the risks the author may not have fully accounted for._
### [Failure path name — e.g., "Capability gap stalls the load-bearing action"]
**The scenario:** [Describe the failure path in concrete terms. What happens, in what order?]
**Likelihood:** [High / Medium / Low — and why.]
**Current mitigation:** [What, if anything, does the strategy do to prevent this? "None" is a valid answer.]
**Suggested mitigation:** [What would reduce the risk.]
### [Next failure path]
...
## Assumption Risk Map
_Load-bearing assumptions ranked by (impact if wrong) x (uncertainty). Focus on the assumptions that could break the strategy, not background assumptions._
| Assumption | Stated in doc? | Impact if wrong | Uncertainty | Risk |
|-----------|----------------|----------------|------------|------|
| [assumption] | Yes / No | High / Medium / Low | High / Medium / Low | [H x H = Critical, etc.] |
| ... | | | | |
[For the top 2-3 highest-risk assumptions, add a paragraph: what happens if this assumption is wrong, and what could the author do now to test it or reduce exposure.]
## Lens Findings
_Include this section when complementary review lenses (7S, Scorecard, Five Forces, Hoshin Kanri) were applied during the review, OR when interview lenses (landscape mapping, choice cascade, value innovation) appear in the strategy document or notes. Omit entirely if neither applies._
**Review lenses applied:** [List which review lenses were activated and why — the trigger condition observed in the document.]
**What the review lenses revealed:**
[For each applied review lens, describe what it found that the core seven dimensions alone would likely have missed. Focus on systemic gaps — the kind that span multiple dimensions or create failure paths the standard review might not catch. 2-4 findings total across all applied lenses.]
- **[Lens name] — [Finding title]**: [What the lens revealed. Point to the specific gap, unstated assumption, or misalignment. Reference the dimension where the finding was also recorded.]
**Interview lens audit:** [Include this subsection only when the strategy document or notes show that Wardley mapping, Playing to Win cascade, or Blue Ocean value innovation analysis was done during the interview.]
[For each interview lens that was applied, assess whether its findings survived into the draft. Did landscape mapping insights sharpen the diagnosis? Did cascade capability gaps appear in assumptions? Did ERRC moves shape the coherent actions? Flag interview lens findings that were dropped or softened — these often represent the sharpest strategic thinking that got polished away.]
- **[Interview lens] — [Assessment]**: [Whether the lens findings are reflected in the draft, partially reflected, or missing. If missing, what was lost.]
## Notes Cross-Reference
_Include this section only when strategy-notes.md or equivalent reasoning notes were available for review._
- **Open questions still open:** [Which questions from the notes remain unresolved in the draft? Are any of them show-stoppers?]
- **Patterns that crept back:** [Bad-strategy patterns caught during the interview that reappeared in the draft, possibly in softer language.]
- **Thinking that was sharpened then softened:** [Places where the notes show a sharper version of the diagnosis or policy than the draft contains.]
- **Lens findings not reflected:** [Landscape mapping, cascade, or value innovation findings from the notes that didn't make it into the draft.]
## Recommended Next Steps
_Ordered by impact. What should the author do with this review?_
1. **[Action]** — [Why this is the highest priority, what it addresses.]
2. **[Action]** — ...
3. **[Action]** — ...---
Notes on producing the review
- Write the review file in the user's working directory unless they specify another location.
- Quote the document. Don't make assertions about what it says — point to the passages. "The diagnosis states: '[quoted text]'" is credible. "The diagnosis is vague" without evidence is not.
- Keep the review proportional to the document. A short strategy memo doesn't need a 2000-word review. Match depth to depth.
- Critical Findings come before Dimension Ratings because they are the highest-value section. A reader who stops after Critical Findings should still walk away with the most important feedback. Dimension Ratings provide the supporting evidence and complete picture.
- The Failure Path Analysis section is where the most unique value lives. Most reviewers tell you what's wrong with the document; this section tells you what could go wrong in the real world because of what's in (or missing from) the document. Invest thought here.
- If the notes cross-reference reveals that the interview produced better thinking than the draft contains, say so directly. Drafts often smooth away the edges that made the thinking sharp.
- After writing, summarize in chat: overall assessment in one sentence, the single most important finding, and the recommended next action. Then stop.