
Generating Novel Ideas
- 61 installs
- 3 repo stars
- Updated June 29, 2026
- tristanmanchester/agent-skills
Helps with ai & agent building tasks during AI-assisted development.
About
generating-novel-ideas is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted coding.
- generating-novel-ideas
- AI & Agent Building
- AI-coding skill
Generating Novel Ideas by the numbers
- 61 all-time installs (skills.sh)
- Ranked #6,381 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/tristanmanchester/agent-skills --skill generating-novel-ideasAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 61 |
|---|---|
| repo stars | ★ 3 |
| Last updated | June 29, 2026 |
| Repository | tristanmanchester/agent-skills ↗ |
What it does
Helps with ai & agent building tasks during AI-assisted development.
Files
Generating Novel Ideas
This skill turns ideation into a search process, not a list-making exercise. The job is to discover a portfolio of distinct, high-potential concepts, not ten polished variations of the first plausible answer.
Critical rules
- Fight collapse. LLMs drift towards fluent sameness. Use independent idea pools before
comparing ideas.
- Prefer concrete mechanisms over vibes. Every finalist needs a sharp twist, an entry
wedge, and a cheap test.
- Separate divergence from judgement. Do not score too early.
- Use ordinary stakeholder or practitioner perspectives when using personas. Do not
imitate celebrity innovators.
- Research late enough to preserve breadth, but early enough to kill obvious
reinventions before the final recommendation.
- Final outputs should usually be a portfolio with spread across mechanism, audience,
and risk, unless the user explicitly asks for a single winner.
Internal roles
Run these roles in sequence. Keep them separate until synthesis.
1. Explorers widen the search space. 2. Critics attack weak, generic, or unrealistic ideas. 3. The synthesiser assembles the final portfolio.
Do not let the critic appear too early. Do not let the synthesiser merge everything into one blurry compromise.
Default workflow
1. Build an opportunity model 2. Partition the search space 3. Generate independent idea pools 4. Run an analogy transfer pass 5. Resolve key contradictions 6. Audit diversity and regenerate missing directions 7. Critique and repair finalists 8. Ground against reality 9. Present a portfolio and experiments
Step 1: Build an opportunity model
Capture the minimum useful brief:
- User goal
- Target user or audience
- Current status quo and what is frustrating, expensive, risky, slow, or emotionally flat
- Hard constraints
- Success criteria
- Available assets, unfair advantages, channels, or capabilities
- Hidden tensions and trade-offs
- What to avoid
When the prompt is sparse, infer reasonable assumptions and state them briefly.
When the user brings an existing idea, do not start by polishing it directly. First extract the underlying job and generate at least two alternative mechanisms.
Step 2: Partition the search space
Choose 3 to 5 independent pools. Pools must differ on at least two axes.
Good axes include:
- stakeholder viewpoint or ordinary persona
- mechanism or value type
- user moment or time horizon
- adoption path or channel
- ambition level
- trust model or ownership model
Examples of useful pool labels:
- frontline operator, zero new habit
- approver or buyer, proof and risk reduction
- novice user, immediate win
- partner or embedded channel
- bold long-shot system shift
Rules:
- Generate each pool as if it has not seen the others.
- Do not compare, deduplicate, or score until all pools are finished.
- Produce 2 to 4 ideas per pool.
- Keep raw ideas short at first: name, one-line concept, primary user, non-obvious move.
This blind partitioning is the main defence against idea collapse.
Step 3: Generate the pools
Inside each pool:
1. Write 2 to 3 fertile reframing questions. 2. Choose two lenses from references/LENSES.md. 3. Generate the first pass. 4. Do a second internal pass and add at least one idea clearly outside the dominant pattern.
Always include one practical lens and one novelty lens.
If the task is complex, breadth comes before depth. Add new mechanism families before expanding any single family.
Step 4: Run an analogy transfer pass
Do not borrow surface style. Borrow mechanism.
1. Abstract the problem into a mechanism, tension, or pattern. 2. Pick 2 to 4 distant domains. 3. Extract what makes those domains work. 4. Map the mechanism back into the problem. 5. Adapt it for the actual constraints and adoption path.
Every strong final set should contain at least one idea born from far analogy, unless the user explicitly wants only safe, incremental options.
For source domains and transfer patterns, use references/LENSES.md.
Step 5: Resolve key contradictions
Write 1 to 3 contradictions at the heart of the task, such as:
- more trust with less friction
- more customisation with less complexity
- more quality with less expert labour
- faster decision-making with lower risk
- more compliance with less manual work
Generate ideas that resolve the contradiction through separation, defaults, staging, guarantees, modularity, reversible commitment, human review only at critical moments, or new ownership boundaries.
For technical, scientific, or engineering prompts, use the structured contradiction method in references/LENSES.md.
Step 6: Audit diversity
Before refinement, check for hidden sameness.
Look for:
- near-duplicates hidden by new wording
- too many ideas using the same mechanism
- too many aimed at the same user moment
- repeated crutches such as AI assistant, dashboard, marketplace, community,
gamification, personalisation, subscription, or platform
- no spread across pragmatic wedge, strategic differentiator, and bold bet
If the set is clustered, regenerate only the missing directions.
When scripts can run and the set is large, optionally use scripts/diversity_audit.py before convergence.
Step 7: Critique and repair finalists
Choose 3 to 6 finalists. For each one, write:
- strongest reason it could work
- smartest sceptic objection
- repair if possible
- kill it if repair makes it generic or unrealistic
Every finalist card should contain:
- Name
- One-sentence pitch
- Who it is for
- Hidden insight or tension
- Imported mechanism or pattern
- Why it is not just the obvious solution
- Entry wedge
- Main risk
- Cheapest disconfirming test
Step 8: Ground against reality
If current market, technical, cultural, or regulatory reality matters and research is available:
- check whether the idea is already common
- identify incumbents or substitutes
- pressure-test feasibility and compliance
- sharpen the why-now and distribution story
- trim false differentiation claims
Do not research so early that the search space collapses into existing categories.
Step 9: Present the result
Default response structure:
1. Working brief and assumptions 2. Opportunity tensions 3. Search partitions used 4. Raw idea families 5. Final portfolio 6. Recommended next move
Use a portfolio, not just a ranking:
- one pragmatic wedge
- one strategic differentiator
- one bold bet
If the user asks for a single winner, still mention the strongest runner-up and the specific reason it lost.
Hard quality bar
No finalist is complete without:
- a clear non-obvious move
- a believable first user and first context
- a path to adoption or distribution
- a cheap test that could disconfirm it
- an explicit line in this format:
This is not just X. The new move is Y.
If Y is vague, decorative, or generic, the concept is not ready.
For the detailed rubric, use references/EVALUATION.md.
Anti-generic rules
- Do not produce a flat list of features around one core mechanism and call it
diversity.
- Do not hide weak ideas behind fluent prose.
- Do not use AI, agent, community, marketplace, dashboard, platform,
personalisation, or gamification as decoration.
- Do not let naming replace concept work.
- Do not overvalue novelty with no adoption path.
- Do not overvalue feasibility when the idea is indistinguishable from existing
practice.
- Prefer specific trade-offs to magical wins on every dimension.
Mode switching
For domain-specific workflows, use references/MODES.md.
Common modes:
- startup or product opportunity
- research hypothesis or scientific idea
- campaign, content, or creative concept
- naming and verbal concept development
- process, service, or operations redesign
Common failure modes
If the outputs feel generic:
- widen the partitions
- add a stronger far-analogy pass
- write sharper contradictions
- regenerate only missing mechanism families
If the outputs feel clever but unusable:
- reduce ambition by one step
- sharpen the first user and first context
- attach a cheaper test and a narrower wedge
If the outputs all sound similar:
- stop scoring
- restart with blind pools from new viewpoints
- avoid celebrity personas and vague mission statements
Examples
Example 1
User says: I need fresh B2B SaaS ideas for compliance teams.
Actions: 1. Build an opportunity model around buyers, blockers, trust, procurement, and audit. 2. Partition by operator, approver, audit trail, and partner channel. 3. Generate blind pools, then run analogy and contradiction passes. 4. Return a portfolio with wedge, differentiator, bold bet, and tests.
Example 2
User says: This startup idea feels generic. Make it genuinely better.
Actions: 1. Extract the underlying job from the current idea. 2. Generate at least two alternative mechanisms before improving the original. 3. Keep only ideas with sharper wedges and clearer tests.
Example 3
User says: Help me come up with novel research directions in battery diagnostics.
Actions: 1. Build tensions, constraints, missing capabilities, and evidence limits. 2. Use scientific mode from references/MODES.md. 3. Add structured contradiction solving and feasibility pressure-testing.
For trigger tests and maintenance checks, use references/VALIDATION.md.
Evaluation and Portfolio Rubric
Use this rubric after divergence, not before. The goal is not just to rank ideas. The goal is to find ideas that are both genuinely fresh and worth advancing.
Fast scoring dimensions
Use a 1 to 5 scale, or High, Medium, Low if speed matters.
Mechanism novelty
Questions:
- Is the mechanism or ownership model actually different?
- Is the twist clear in one line?
- Would a smart person in the space admit it is directionally fresh?
User value clarity
Questions:
- Is there a clear user, buyer, or beneficiary?
- Does it solve a real pain, create a real gain, or unlock a real opportunity?
- Is the benefit legible without a long explanation?
Feasibility
Questions:
- Is there a plausible first version?
- Can it be tested without heroic resources?
- Does it fit the stated constraints?
Wedge strength
Questions:
- Is there a believable first context?
- Is the first use case narrower and stronger than the long-term vision?
- Is there a route to adoption?
Differentiation
Questions:
- Does it avoid category clichés?
- Is it hard to confuse with the standard answer?
- Does the twist produce a concrete advantage?
Learning velocity
Questions:
- Can the core assumption be tested quickly?
- Is there a cheap disconfirming experiment?
- Would the next step teach something decisive?
Emotional or narrative pull
Questions:
- Does the idea create relief, delight, confidence, status, or intrigue?
- Will people remember the story?
- Is there a simple way to explain why it matters now?
Portfolio audit
A strong final set is not just three high scorers. It is a spread.
Check diversity across:
- mechanism
- user moment
- channel or route to market
- risk profile
- time horizon
- ambition level
Good default portfolio:
- one pragmatic wedge
- one strategic differentiator
- one bold bet
If all finalists use the same engine with different dressing, the portfolio failed.
Kill criteria
Downgrade or kill ideas that:
- depend on magical behaviour change
- require unrealistic trust, data, budget, or distribution
- sound smart but lack a first use case
- are fresh only at the naming layer
- collapse into a common category when explained plainly
- cannot survive the sceptic objection without becoming generic
The twist test
For every finalist, write:
This is not just X. The new move is Y.
If Y is vague, decorative, or interchangeable, the idea is not ready.
The first-use-case test
Every finalist should name:
- first user
- first problem
- first context
- first version
If one of these is missing, the concept is still too abstract.
The cheap-test menu
Attach a low-cost test to each finalist.
Useful options:
- concierge test
- landing-page test
- fake-door test
- Wizard-of-Oz test
- customer interview pack
- prototype sprint
- workflow pilot
- pre-sale or waitlist
- manual simulation of the service
- objection test with real buyers or users
Choose the cheapest test that could disconfirm the idea fastest.
Recommendation format
When giving a final recommendation, include:
- why this idea won
- which assumption matters most
- what to test next
- what evidence would make you double down
- what evidence would make you abandon it
Example finalist card
- Name:
- One-line concept:
- Who it is for:
- Hidden tension:
- Imported mechanism:
- Twist:
- Entry wedge:
- Main risk:
- Cheapest test:
Final caution
Do not end with the safest idea by default. Do not end with the weirdest idea by default. End with the strongest balance of freshness, utility, wedge strength, and learning velocity.
Lenses and Search Patterns
Use 2 to 4 lenses per pool, not all of them. A strong mix usually includes:
- one practical lens
- one novelty lens
- one adoption or trust lens
- one contradiction or simplification lens
Search partitions first
Before choosing lenses, decide how the search space will be partitioned.
Useful partition axes:
- stakeholder or ordinary persona
- user moment
- value type
- trust model
- channel or distribution path
- ownership model
- time horizon
- ambition level
Example partition sets:
- operator, approver, auditor, partner
- before the job, during the risky moment, after the outcome
- save time, reduce risk, increase status, create access
- lightweight wedge, adjacent expansion, bold system shift
Do not let every pool share the same mechanism.
Reframing patterns
Before generating ideas, write 2 to 3 fertile questions inside each pool.
Useful patterns:
- How might we create value before commitment?
- How might we remove the scariest moment?
- How might we make the risky choice feel safe?
- How might we make the safe choice feel more valuable?
- How might we turn the hardest constraint into the advantage?
- How might we move the value earlier, later, smaller, or more reversible?
- How might we solve the opposite problem on purpose and learn from it?
- How might we help the user without asking for a new habit?
Lens 1: Far analogy
Use when the space feels saturated.
Process:
1. Abstract the problem to its mechanism or tension. 2. Pick a distant domain. 3. Extract the mechanism, not the surface style. 4. Rebuild the mechanism in the target context.
Good distant domains:
- logistics networks
- insurance
- amateur sports
- luxury hospitality
- museums
- tax preparation
- hospitals
- multiplayer games
- airports
- wedding planning
- theme parks
- disaster response
Good mechanism transfers:
- queue management
- reassurance and guarantees
- progressive disclosure
- expert triage
- shared rituals
- checkpoints and hand-offs
- replay and review
- status signalling
- reversible commitment
- exception handling
Lens 2: Ordinary persona or stakeholder
Use when ideas are converging too early.
Prefer grounded perspectives such as:
- junior analyst
- frontline operator
- team lead
- compliance approver
- procurement reviewer
- channel partner
- sceptical customer
- occasional user
- support agent
Do not imitate heroic or celebrity personas. The goal is knowledge partitioning, not style mimicry.
For each persona, ask:
- What do they see that others do not?
- What do they fear?
- What do they measure?
- What looks expensive, risky, or annoying from their seat?
- What would count as an immediate win?
Lens 3: Structured contradiction solving
Use for engineering, science, operations, and stubborn product problems.
1. Write the contradiction clearly. 2. Ask whether the solution should be separated by time, place, user, interface, confidence level, or ownership. 3. Try one of these patterns:
- defaults first, control later
- manual at the risky moment, automated elsewhere
- reversible commitment
- guarantee or escrow
- modularity
- staged disclosure
- simulation before action
- sampling before full purchase
- separate novice and expert modes
- shift the burden to a different actor
- sell certainty rather than features
- transform an exception into the product
The point is to stop treating trade-offs as fixed.
Lens 4: Constraint-first
Use when the limits are real.
Prompts:
- If this had to work in one week, what is the wedge?
- If this needed zero new behaviour, what survives?
- If legal or procurement had to approve it, what changes?
- If data were limited, what would still be valuable?
- If this had to work in low-trust settings, what emerges?
- If the budget were halved, what new model appears?
Constraints are not just filters. They often reveal stronger concepts than blank-page brainstorming.
Lens 5: Distribution wedge
Use when the concept needs a believable route to users.
Ask:
- Where does the user already work or gather?
- What adjacent workflow could carry the idea in?
- What partner or integration makes adoption easier?
- What small result would make people share it?
- What single use case is strong enough to earn distribution?
A slightly less novel idea with a sharp wedge can beat a clever idea with no path in.
Lens 6: Incentives and trust
Use when fear, proof, reputation, or conflicting stakeholders shape adoption.
Ask:
- Who benefits?
- Who hesitates?
- Who blocks?
- What proof would unlock action?
- What guarantee, audit trail, or reversible choice changes behaviour?
- What would make the risky option feel safe?
- What would make the safe option feel worthwhile?
Lens 7: Time-shift
Use when the pain is concentrated around a moment.
Prompts:
- What should happen before the pain appears?
- What should happen during the risky moment?
- What should happen after the outcome?
- What could be persistent instead of episodic?
- What could be reversible instead of final?
Lens 8: Business model or ownership flip
Use when the current payment or packaging logic blocks adoption.
Prompts:
- What changes if someone else pays?
- What changes if payment happens after value?
- What changes if the user buys certainty instead of features?
- What changes if the product becomes a service, audit, benchmark, guarantee, or
outcome?
- What changes if ownership moves to the team, partner, or platform?
Lens 9: Identity, emotion, and status
Use for consumer, education, team tools, creative products, or any emotionally loaded decision.
Ask:
- What identity is the user trying to protect or project?
- What emotion dominates the moment?
- What would make the user feel competent, seen, respected, or ahead?
- What part of the experience feels cold, invisible, or thankless?
Lens 10: Extreme simplification
Use when the category is bloated.
Prompts:
- What is the smallest version that still feels magical?
- What single output or action could stand in for the whole system?
- What can be removed so the value is obvious in seconds?
- What focused version would be memorable because it does less?
Lens 11: Hybridisation
Use after the first pass, not before.
1. Pick two ideas with different strengths. 2. Name the useful property from each. 3. Combine only those properties. 4. Drop any complexity that does not strengthen the wedge.
Good hybrids become clearer. Bad hybrids become crowded.
Mode suggestions
Startup or product opportunity
Start with:
- far analogy
- constraint-first
- distribution wedge
- business model or ownership flip
- incentives and trust
Research hypothesis or scientific idea
Start with:
- structured contradiction solving
- far analogy from adjacent instruments or fields
- time-shift
- constraint-first
- extreme simplification
Campaign or creative concept
Start with:
- identity, emotion, and status
- far analogy
- time-shift
- distribution wedge
- hybridisation
Process, service, or operations redesign
Start with:
- structured contradiction solving
- incentives and trust
- time-shift
- constraint-first
- business model or ownership flip
Domain Modes
Use these mode-specific adjustments when the task has a clear domain. The main workflow still applies.
Startup or product opportunity
Priorities:
- hidden user tension
- route to adoption
- why-now signal
- business model or ownership logic
- wedge first, platform later
Extra questions:
- What adjacent workflow can carry this in?
- Who signs off or blocks adoption?
- What is the smallest offer that is still compelling?
- What evidence would prove this is not just another feature idea?
Best final format:
- pragmatic wedge
- strategic differentiator
- bold bet
- tests and adoption assumptions
Research hypothesis or scientific idea
Priorities:
- novelty with plausible mechanism
- evidence limits
- data or instrumentation requirement
- experimental tractability
- failure conditions
Extra questions:
- What assumption or theory is being challenged?
- What observation would falsify it?
- What instrument, dataset, or proxy could test it cheaply?
- What makes it different from a standard literature extrapolation?
Best final format:
- hypothesis
- why it is interesting
- why it might work
- key uncertainty
- first experiment
- failure signal
Avoid purely magical proposals with no measurement path.
Campaign, content, or creative concept
Priorities:
- cultural tension
- emotional hook
- memorable framing
- spread mechanism
- assetability across formats
Extra questions:
- What identity or feeling is being activated?
- What visual or verbal hook carries the territory?
- What makes it shareable or discussable?
- Can one core idea generate many executions without going thin?
Best final format:
- concept territory
- central tension
- hook
- sample executions
- channel fit
- risk
Naming and verbal concept development
Naming is downstream of concept design.
Workflow:
1. Create 3 to 5 concept territories first. 2. Pick one or two territories. 3. Generate names within those territories. 4. Explain the logic behind the strongest names.
Useful name directions:
- literal and clear
- evocative
- metaphorical
- contrast-based
- status-based
- coined or compressed
- procedural or active
Do not start with name lists before the concept families exist.
Practical caution:
- flag that names still need legal and market checks
- do not claim clearance without actual verification
Process, service, or operations redesign
Priorities:
- bottlenecks
- hand-offs
- approval and exception paths
- trust and reversibility
- who carries the burden
Extra questions:
- Where does work wait?
- Where does knowledge get lost?
- Which exception cases consume the team?
- What proof or audit trail would unblock action?
- What step should be manual only when risk is high?
Best final format:
- redesigned flow
- what changes
- why it is better
- implementation wedge
- operational risk
- pilot
Validation and Trigger Tests
Use this file to maintain the skill over time.
Should trigger
These prompts should strongly activate the skill:
- Help me brainstorm differentiated app ideas for tradespeople.
- This product concept is boring. Find genuinely fresher directions.
- I need novel research directions in battery diagnostics.
- Give me campaign territories for a new coffee brand.
- We need creative but commercially sensible membership ideas for a museum.
- Name this product, but first strengthen the concept.
- Find non-obvious service ideas for an imaging consultancy.
- Help us escape generic feature brainstorming for our SaaS roadmap.
Should also trigger on paraphrases
- Come up with fresh options
- Ideate around this
- Push this into more original territory
- Break me out of generic answers
- Find stronger concept families
- What are some non-obvious angles here?
Should not trigger
These prompts should usually not activate the skill:
- Proofread this email.
- Summarise this report.
- What is the capital of Peru?
- Give me the latest semiconductor news.
- Translate this paragraph into German.
- Write the code for this exact API spec.
Edge case:
- If the user asks to design options, explore alternatives, or invent stronger concepts
before implementation, the skill should trigger even if code or documents are part of the eventual output.
Quick quality checks
A healthy run should usually show:
- at least three distinct mechanism families in the raw set
- at least one far-analogy-derived concept when novelty matters
- at least one clear adoption wedge in the finalists
- one cheap test attached to every finalist
- a final portfolio with spread across risk and ambition
Failure signals
The skill is underperforming if:
- the final ideas are polished but obviously similar
- every concept uses the same category cliché
- no idea has a believable first user and first context
- the final recommendation could have been produced by a generic brainstorming prompt
- the naming layer is doing all the work
Manual regression suite
Run these tests after edits:
1. A blank-page startup prompt 2. A weak existing concept that needs improvement 3. A research idea prompt 4. A campaign or creative territory prompt 5. A process redesign prompt
For each run, check:
- Did the skill partition the search space?
- Did it delay judgement?
- Did it use at least one novelty lens and one practical lens?
- Did it criticise and repair finalists?
- Did it present a portfolio rather than near-duplicates?
\
#!/usr/bin/env python3
"""
Audit raw idea sets for near-duplicates and dominant patterns.
Accepted input:
- JSON list of strings
- JSON list of objects with name and concept or description
- Plain text with one idea per line
- Markdown bullets
Usage:
python scripts/diversity_audit.py ideas.json
python scripts/diversity_audit.py ideas.txt
cat ideas.txt | python scripts/diversity_audit.py
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from collections import Counter
from dataclasses import dataclass
from difflib import SequenceMatcher
from pathlib import Path
from typing import Iterable, List
STOPWORDS = {
"a", "an", "and", "are", "as", "at", "be", "but", "by", "for", "from", "how",
"if", "in", "into", "is", "it", "its", "of", "on", "or", "our", "that", "the",
"their", "there", "this", "to", "using", "with", "we", "you", "your", "than",
"then", "will", "can", "could", "should", "would", "make", "help", "idea",
"ideas", "new", "better"
}
TAG_PATTERNS = {
"ai_or_agent": [r"\bai\b", r"\bagent\b", r"\bassistant\b", r"\bchatbot\b", r"\bllm\b"],
"dashboard": [r"\bdashboard\b", r"\banalytics\b", r"\breport\b", r"\bmonitoring\b"],
"marketplace": [r"\bmarketplace\b", r"\bplatform\b", r"\bnetwork\b"],
"community": [r"\bcommunity\b", r"\bforum\b", r"\bmember\b"],
"gamification": [r"\bgamif", r"\bpoints\b", r"\bbadge\b", r"\bleaderboard\b"],
"personalisation": [r"\bpersonali", r"\brecommend", r"\bcustomi", r"\btailor"],
"automation": [r"\bautomation\b", r"\bautomate\b", r"\bworkflow\b", r"\borchestrat"],
"trust": [r"\bguarantee\b", r"\bproof\b", r"\baud(it|it trail)\b", r"\bescrow\b", r"\breversible\b"],
"service": [r"\bconcierge\b", r"\bservice\b", r"\bmanual\b", r"\bhuman\b"],
"business_model": [r"\bsubscription\b", r"\bpricing\b", r"\boutcome\b", r"\bpay\b"],
}
VALUE_PATTERNS = {
"time": [r"\bfaster\b", r"\bsave time\b", r"\bminutes\b", r"\bworkflow\b"],
"risk": [r"\brisk\b", r"\bcompliance\b", r"\btrust\b", r"\bsafe\b", r"\bguarantee\b"],
"money": [r"\bcost\b", r"\brevenue\b", r"\bprice\b", r"\bpay\b"],
"status": [r"\bstatus\b", r"\breputation\b", r"\bprestige\b"],
"access": [r"\baccess\b", r"\bavailability\b", r"\bdistribution\b", r"\bchannel\b"],
"delight": [r"\bdelight\b", r"\bfun\b", r"\bjoy\b", r"\bmagic\b"],
}
CHANNEL_PATTERNS = {
"embedded": [r"\bintegration\b", r"\bplugin\b", r"\bembedded\b", r"\bin app\b"],
"partner": [r"\bpartner\b", r"\bchannel\b", r"\breseller\b"],
"content_or_referral": [r"\bshare\b", r"\breferral\b", r"\bcontent\b", r"\bviral\b"],
"sales_led": [r"\bprocurement\b", r"\benterprise\b", r"\bsales\b"],
"self_serve": [r"\bself serve\b", r"\bsign up\b", r"\bfree trial\b", r"\bwaitlist\b"],
}
@dataclass
class Idea:
idx: int
name: str
text: str
@property
def combined(self) -> str:
return f"{self.name}. {self.text}".strip(". ")
def load_text_from_path_or_stdin(path: str | None) -> str:
if path:
return Path(path).read_text(encoding="utf-8")
if sys.stdin.isatty():
raise SystemExit("Provide a file path or pipe input on stdin.")
return sys.stdin.read()
def parse_ideas(raw: str) -> List[Idea]:
stripped = raw.strip()
if not stripped:
raise SystemExit("No input ideas found.")
# Try JSON first.
try:
data = json.loads(stripped)
if isinstance(data, list):
ideas: List[Idea] = []
for i, item in enumerate(data, start=1):
if isinstance(item, str):
ideas.append(Idea(i, f"Idea {i}", item.strip()))
elif isinstance(item, dict):
name = str(item.get("name") or item.get("title") or f"Idea {i}").strip()
text = str(
item.get("concept")
or item.get("description")
or item.get("idea")
or item.get("summary")
or ""
).strip()
if not text:
text = json.dumps(item, ensure_ascii=False)
ideas.append(Idea(i, name, text))
else:
ideas.append(Idea(i, f"Idea {i}", str(item).strip()))
return [idea for idea in ideas if idea.text]
except json.JSONDecodeError:
pass
lines = []
for line in stripped.splitlines():
cleaned = re.sub(r"^\s*[-*+]\s+", "", line).strip()
cleaned = re.sub(r"^\s*\d+[.)]\s+", "", cleaned).strip()
if cleaned:
lines.append(cleaned)
ideas = []
for i, line in enumerate(lines, start=1):
if ":" in line and len(line.split(":", 1)[0].split()) <= 6:
name, text = line.split(":", 1)
ideas.append(Idea(i, name.strip(), text.strip()))
else:
ideas.append(Idea(i, f"Idea {i}", line))
if not ideas:
raise SystemExit("Could not parse any ideas from the input.")
return ideas
def normalise_tokens(text: str) -> List[str]:
words = re.findall(r"[a-z0-9]+", text.lower())
return [w for w in words if w not in STOPWORDS and len(w) > 2]
def jaccard_similarity(a: Iterable[str], b: Iterable[str]) -> float:
set_a = set(a)
set_b = set(b)
if not set_a or not set_b:
return 0.0
return len(set_a & set_b) / len(set_a | set_b)
def sequence_similarity(a: str, b: str) -> float:
return SequenceMatcher(None, a.lower(), b.lower()).ratio()
def apply_patterns(text: str, patterns: dict[str, list[str]]) -> Counter:
counts: Counter = Counter()
for tag, regs in patterns.items():
if any(re.search(reg, text, flags=re.IGNORECASE) for reg in regs):
counts[tag] += 1
return counts
def find_duplicates(ideas: List[Idea]) -> list[tuple[Idea, Idea, float, float]]:
results = []
token_map = {idea.idx: normalise_tokens(idea.combined) for idea in ideas}
for i, a in enumerate(ideas):
for b in ideas[i + 1:]:
jac = jaccard_similarity(token_map[a.idx], token_map[b.idx])
seq = sequence_similarity(a.combined, b.combined)
if jac >= 0.33 or seq >= 0.72:
results.append((a, b, jac, seq))
return sorted(results, key=lambda x: max(x[2], x[3]), reverse=True)
def top_terms(ideas: List[Idea], limit: int = 12) -> list[tuple[str, int]]:
counter = Counter()
for idea in ideas:
counter.update(normalise_tokens(idea.combined))
return counter.most_common(limit)
def coverage_counts(ideas: List[Idea], patterns: dict[str, list[str]]) -> Counter:
counter = Counter()
for idea in ideas:
counter.update(apply_patterns(idea.combined, patterns))
return counter
def suggest_regeneration(
tag_counts: Counter,
value_counts: Counter,
channel_counts: Counter,
total: int,
) -> list[str]:
suggestions: list[str] = []
dominant_tags = [tag for tag, count in tag_counts.items() if count / max(total, 1) >= 0.4]
if dominant_tags:
joined = ", ".join(dominant_tags)
suggestions.append(
f"Generate 3 ideas that do not use these dominant patterns: {joined}."
)
if value_counts.get("risk", 0) == 0:
suggestions.append("Generate 2 ideas where the main value is risk reduction or reassurance.")
if value_counts.get("time", 0) == 0:
suggestions.append("Generate 2 ideas where the main value is time compression or effort removal.")
if value_counts.get("status", 0) == 0:
suggestions.append("Generate 1 or 2 ideas built around status, recognition, or visible competence.")
if value_counts.get("access", 0) == 0:
suggestions.append("Generate 1 or 2 ideas that win through access, channel, or distribution.")
if sum(channel_counts.values()) == 0:
suggestions.append("Generate 2 ideas whose wedge is a partner, embedded workflow, or referral channel.")
if not suggestions:
suggestions.append("Generate 2 ideas from far analogies in logistics, insurance, or museums.")
suggestions.append("Generate 2 ideas that change ownership, payment timing, or who carries the risk.")
return suggestions
def format_markdown(ideas: List[Idea]) -> str:
duplicate_pairs = find_duplicates(ideas)
tag_counts = coverage_counts(ideas, TAG_PATTERNS)
value_counts = coverage_counts(ideas, VALUE_PATTERNS)
channel_counts = coverage_counts(ideas, CHANNEL_PATTERNS)
term_counts = top_terms(ideas)
suggestions = suggest_regeneration(tag_counts, value_counts, channel_counts, len(ideas))
lines: list[str] = []
lines.append("# Diversity Audit")
lines.append("")
lines.append(f"Ideas analysed: {len(ideas)}")
lines.append("")
lines.append("## Likely near-duplicates")
if duplicate_pairs:
for a, b, jac, seq in duplicate_pairs[:10]:
lines.append(
f"- {a.idx} and {b.idx}: Jaccard {jac:.2f}, sequence {seq:.2f} "
f"— {a.name} / {b.name}"
)
else:
lines.append("- No obvious near-duplicates detected by the heuristic.")
lines.append("")
lines.append("## Dominant repeated patterns")
if tag_counts:
for tag, count in tag_counts.most_common():
lines.append(f"- {tag}: {count}")
else:
lines.append("- No repeated cliché patterns detected.")
lines.append("")
lines.append("## Coverage signals")
if value_counts:
lines.append("- Value types:")
for tag, count in value_counts.most_common():
lines.append(f" - {tag}: {count}")
else:
lines.append("- Value types: no strong signals detected")
if channel_counts:
lines.append("- Channel signals:")
for tag, count in channel_counts.most_common():
lines.append(f" - {tag}: {count}")
else:
lines.append("- Channel signals: no clear wedge language detected")
lines.append("")
lines.append("## Repeated terms")
if term_counts:
lines.append("- " + ", ".join(f"{term} {count}" for term, count in term_counts))
else:
lines.append("- No repeated terms detected.")
lines.append("")
lines.append("## Suggested regeneration prompts")
for item in suggestions:
lines.append(f"- {item}")
lines.append("")
lines.append("## Notes")
lines.append(
"- This is a lightweight heuristic audit. Use it to spot collapse, not to replace judgement."
)
lines.append(
"- If the audit flags many duplicates, regenerate missing mechanism families before scoring."
)
return "\n".join(lines)
def main() -> None:
parser = argparse.ArgumentParser(description="Audit idea lists for near-duplicates and dominant patterns.")
parser.add_argument("path", nargs="?", help="Path to a JSON or text file containing ideas.")
args = parser.parse_args()
raw = load_text_from_path_or_stdin(args.path)
ideas = parse_ideas(raw)
report = format_markdown(ideas)
print(report)
if __name__ == "__main__":
main()