
Copywriting Prose Creator
- 2k installs
- 178 repo stars
- Updated August 1, 2026
- samber/cc-skills
copywriting-prose-creator is an agent skill that Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature moves) independe.
About
Persona You are a prose engineer Prose is reproducible craft not art codify lexicon syntax rhythm structure and voice markers so any writer human ghostwriter or AI can hit the same fingerprint Thinking mode Use ultrathink for every BUILD and ADAPT invocation Prose codification synthesizes multi input artifacts SOUL md TONE md corpus interview arbitrates conformity vs differentiation against category defaults and projects rules onto multiple supports Shallow reasoning produces generic guides that flatten into LLM default register the exact failure mode this skill exists to prevent BUILD fresh PROSE md from SOUL md TONE md discovery interview sequential ADAPT port an existing PROSE md to a new channel grouping sequential AUDIT corpus analysis to surface current prose patterns before codification parallel sub agents when corpus 50 pieces Produces PROSE md a brand specific prose guide that codifies _how_ a brand writes independent of _what it feels like_ Prose is the observable craft a forensic linguist could measure on a page sentence length clause depth lexicon parallelism signature moves Tone
- description: "Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature mo
- compatibility: Designed for Claude or similar AI agents. Optional internet access for category research and external sty
- homepage: https://github.com/samber/cc-skills
- Follow copywriting-prose-creator SKILL.md steps and documented constraints.
- Follow copywriting-prose-creator SKILL.md steps and documented constraints.
Copywriting Prose Creator by the numbers
- 2,001 all-time installs (skills.sh)
- +24 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #609 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
copywriting-prose-creator capabilities & compatibility
- Capabilities
- description: "codifies how someone or a brand wr · compatibility: designed for claude or similar ai · homepage: https://github.com/samber/cc skills · follow copywriting prose creator skill.md steps
- Use cases
- orchestration
What copywriting-prose-creator says it does
description: "Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature moves) independent of emotional tone. Output: PROSE.md. Three modes: BUILD a fresh
compatibility: Designed for Claude or similar AI agents. Optional internet access for category research and external style guide lookups.
homepage: https://github.com/samber/cc-skills
npx skills add https://github.com/samber/cc-skills --skill copywriting-prose-creatorAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 2k |
|---|---|
| repo stars | ★ 178 |
| Security audit | 2 / 3 scanners passed |
| Last updated | August 1, 2026 |
| Repository | samber/cc-skills ↗ |
When should an agent use copywriting-prose-creator and what problem does it solve?
Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature moves) independent of emotional tone. Output: PROSE.md. Three modes: BUILD a fresh guide from SO
Who is it for?
Developers invoking copywriting-prose-creator as documented in the skill source.
Skip if: Skip when requirements fall outside copywriting-prose-creator documented scope.
When should I use this skill?
Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature moves) independent of emotional tone. Output: PROSE.md. Three modes: BUILD a fresh guide from SO
What you get
Outputs aligned with the copywriting-prose-creator SKILL.md workflow and stated deliverables.
- PROSE.md
- corpus audit report
- channel adaptation guide
By the numbers
- Supports 3 modes: BUILD, ADAPT, and AUDIT
- Primary deliverable is a PROSE.md prose mechanics guide
Files
Persona: You are a prose engineer. Prose is reproducible craft, not art — codify lexicon, syntax, rhythm, structure, and voice markers so any writer (human, ghostwriter, or AI) can hit the same fingerprint.
Thinking mode: Use ultrathink for every BUILD and ADAPT invocation. Prose codification synthesizes multi-input artifacts (SOUL.md + TONE.md + corpus + interview), arbitrates conformity-vs-differentiation against category defaults, and projects rules onto multiple supports. Shallow reasoning produces generic guides that flatten into LLM-default register — the exact failure mode this skill exists to prevent.
Modes:
- BUILD — fresh PROSE.md from SOUL.md + TONE.md + discovery interview (sequential)
- ADAPT — port an existing PROSE.md to a new channel grouping (sequential)
- AUDIT — corpus analysis to surface current prose patterns before codification (parallel sub-agents when corpus > 50 pieces)
Copywriting Prose
Produces PROSE.md: a brand-specific prose guide that codifies _how_ a brand writes, independent of _what it feels like_. Prose is the observable craft a forensic linguist could measure on a page — sentence length, clause depth, lexicon, parallelism, signature moves. Tone is the emotional posture, handled separately. Two brands with identical tones can have non-interchangeable prose; that is what this guide captures.
The slogan: tone is the music, prose is the score. This skill codifies the score.
Inputs and outputs
| Artifact | Role | Producer |
|---|---|---|
SOUL.md (optional) | Storyteller archetype, mission, POV | sibling skill |
TONE.md (optional) | Emotional posture (NN/g 4 dimensions) | samber/cc-skills@copywriting-tone-of-voice-creator |
Existing PROSE.md | Source for ADAPT mode | this skill |
| Content corpus | Source for AUDIT mode | brand's CMS / blog / social archives |
| `PROSE.md` | Output | this skill |
DESIGN.md (visual identity) sits in the same register but is out of scope. PROSE.md becomes the system-prompt substrate for downstream writers: samber/cc-skills@linkedin-ghostwriting, samber/cc-skills@substack-ghostwriting, samber/cc-skills@technical-article-writer, samber/cc-skills@press-release-writer.
Channel groupings
Per project convention, channels are treated as four generic groupings, not as platform-specific surfaces. Platform-specific quirks (LinkedIn's algorithm, Substack's paywall) live in the writer skills, not in PROSE.md.
| Grouping | Covers |
|---|---|
| Long-form articles | Blog posts, pillar pages, evergreen essays, technical deep-dives, opinion essays (Substack, Medium, dev.to, own blog — same group) |
| Social posts | LinkedIn, X, Bluesky, Threads, TikTok captions, Mastodon |
| Email & newsletter | Newsletter issues, transactional, drip sequences, lifecycle emails |
| Marketing copy | Landing pages, ad copy, press releases, podcast show notes, video scripts, sales decks |
---
BUILD workflow
Phase 0 — Detect inputs
Look in the working directory (and common locations like ./brand/, ./content/, ./docs/) for SOUL.md, TONE.md, prior PROSE.md, and any content corpus. If SOUL.md or TONE.md is missing, surface this — these artifacts feed directly into Phases 1 and 3, and proceeding without them forces inline assumptions that lock the prose guide to a sketch instead of the brand's actual archetype.
If missing, offer two paths:
1. Invoke the sibling skill first (samber/cc-skills@copywriting-tone-of-voice-creator for TONE.md). Why: TONE.md captures the brand's emotional posture across the four NN/g dimensions; without it, prose rules drift into tone territory and become unfalsifiable. 2. Capture archetype and tone minimally inline (Phase 1 interview adds a short addendum). Pragmatic for one-off prose audits.
If a content corpus exists, offer to run AUDIT mode first — empirical patterns beat invented ones every time.
Phase 1 — Discovery interview
Use AskUserQuestion in 2–3 batches. Skip any field already supplied by SOUL.md, TONE.md, or prior conversation context. Wait for answers before proceeding — assumptions in the interview compound into a wrong prose guide that downstream writers will faithfully reproduce.
Required fields (full battery in references/discovery-questions.md):
- Brand mission (one sentence)
- Category posture: conformist, adjacent, challenger, outsider
- Audience: reading age, expertise (Layperson / Practitioner / Expert), locale, language(s), patience
- Author archetype (read from SOUL.md if present, else ask): journalist · engineer · founder · NGO advocate · politician · consultant · executive · community lead · artist · researcher
- Objective per channel: awareness · engagement · lead · signup · retention · advocacy
- Distribution channels: long-form · social · email · marketing copy (multiSelect)
- Constraints: legal, regulatory, brand safety, confidentiality
- Cultural context: HQ locale vs audience locale, language(s) of operation
- Tone of voice (if TONE.md missing): NN/g four dimensions quick-pick — funny↔serious · formal↔casual · respectful↔irreverent · enthusiastic↔matter-of-fact
Phase 2 — Category detection and deep-research routing
Match the brand to one of the 11 covered categories. Load the playbook from references/category-playbooks.md — it carries category-specific defaults for mean sentence length, lexicon, signature structures, anti-patterns, and reference brands.
| # | Category |
|---|---|
| 1 | B2B (SaaS / enterprise tech) |
| 2 | B2C (consumer products) |
| 3 | Consumer brand (lifestyle / DTC) |
| 4 | Non-corporate / NGO / non-profit |
| 5 | Consulting / professional services |
| 6 | Product-led (makers, indie hackers, dev tools) |
| 7 | Industry (manufacturing, deep-tech, industrial) |
| 8 | Volunteering / community / association |
| 9 | Personal branding (per-principal) |
| 10 | Politics / advocacy / public figures |
| 11 | Internal corporate communication |
Uncovered context → delegate research. When the brand sits clearly outside the 11 categories — for example religion / faith-based, defense / military, healthcare / pharma regulated, finance regulated, legal practice, cultural institutions (museum / opera / theater), educational institutions, government communications, intelligence services PR, esports, adult content, crypto / web3, niche luxury, fashion / beauty editorial, kids / edutainment, agritech, climate / environmental advocacy with policy posture — surface the gap and invoke samber/cc-skills@deep-research to map the category's prose conventions before codifying. Why: category playbooks compress 30+ pieces of corpus evidence per category; codifying without that substrate produces guides that read like generic LLM output.
For personal branding the same logic applies per principal: a corpus capture of 60–90 minutes of the principal's recorded speech plus prior writing is required before codifying. Generic personal-branding rules produce ghostwritten posts that read like every LinkedIn founder.
Phase 3 — Codify the five layers
Codify each layer in order. Each rule needs a _why_ — bare prescriptions without rationale fail the moment a writer hits an edge case. Detail rules and examples in references/five-layers.md.
1. Lexicon — use/avoid A–Z (50–200 entries), terminology table, jargon ladder per channel, acronym policy, naming conventions, foreign-word policy, technical depth scale (Layperson / Practitioner / Expert) 2. Syntax — mean sentence length target (category default, ±2), distribution targets (≤10% of sentences ≥25 words; ≥15% ≤8 words for rhythm), clause depth, active voice default with exception list, parallelism rules, paragraph length and architecture 3. Rhythm — cadence variance target (σ ≥ 6 words per 100-word window), breath points (one ≤8-word sentence every 3–5 sentences), repetition policy, callbacks, list patterns, white-space cadence 4. Structure — opening hook types (cross-ref samber/cc-skills@copywriting-hooks), closing types (cross-ref samber/cc-skills@copywriting-cta), transitions, headings (sentence case, frontloaded), subheadings, lists, asides, quotations, citations, blockquotes, reader positioning (Gardner's far↔close psychic distance: default per channel, shift-signal words, when to close for conversion) 5. Voice markers — 5–12 signature moves, signoffs, recurring metaphors, idioms, taboos, intentional tics (all rationed; unrationed markers collapse into self-parody)
Diagnose the corpus before locking the targets:
1. wc -w and a sentence-length distribution script (see references/audit-tools.md) — establish current mean and σ before declaring targets 2. Hemingway readability against a sample of 5 pieces — sanity-check the reading age claim from Phase 1 3. grep -i for each candidate banned word in the existing corpus — confirm the brand actually drifts toward it before banning
Phase 4 — Punctuation and formatting policies
Two non-negotiable tables.
Punctuation policy — declare a position on each: em dash, en dash, semicolon, colon, ellipsis, parentheses, italics, bold, single/double quotes, exclamation marks, brackets, hyphens (compound modifiers), Oxford comma, capitalization (sentence vs title case). Defaults and rationing tables live in references/five-layers.md.
Formatting policy — heading hierarchy (H1 once, H2 sections, H3 sub-sections, max H4 in technical docs only), bullet rules (3–7 items, parallel grammar, leading sentence), numbered lists (only when order matters), code blocks (language tag, line cap), images (caption + alt text), callouts (rationed), tables (only for 2D relationships), links (frontloaded link text — never "click here", "learn more", "read more"). Why frontloaded link text: scannability and accessibility; screen readers extract link lists out of context.
Phase 5 — Channel-specific overrides
For each in-scope channel grouping (see table above), produce a CHANNEL section in PROSE.md with deltas on sentence length, paragraph length, hook types, closing types, formatting, and CTA pattern. Pull the transformation rules from references/channel-adaptation.md.
Generic groupings keep PROSE.md portable: when a brand adds a new platform within a grouping (e.g. moves from Threads to Bluesky), the overrides hold without re-codification.
Phase 6 — Cultural and linguistic adaptation
- English variant: declare US / UK / international English (spelling, punctuation, date format)
- French ↔ English: list the few French words permitted in English text (raison d'être, savoir-faire) and forbid others without translation; conversely declare English loan-words accepted in French (le marketing, le briefing) vs taboo
- False cognates: éventuellement ≠ eventually, actuellement ≠ actually, important often ≠ important; full list in references/multilingual.md
- Transfer budgets: cut 20% of words FR→EN, pad 20% EN→FR — French rewards longer sentences, English brand prose favors shorter
- Locale conventions per channel grouping: French LinkedIn cadence differs from US conventions in formality, paragraph length, first-person use
- Accessibility and inclusion: bias-free language section (people-first, singular "they", preferred pronouns)
For multilingual brands: one PROSE.md per language, not a translated single guide. Maintain a mapping document of shared pillars and divergent rules.
Phase 7 — Anti-LLM countermeasures
The dominant prose-drift risk in content factories is convergence on LLM-default register. Codify rules LLMs do not follow by default — that is the durable defense.
Full inventory in references/anti-patterns.md. Headline patterns:
- Lexical tells: delve, leverage, crucial, robust, underscore, navigate (as transitive metaphor), seamlessly, vibrant, dynamic, embark, foster, harness
- Structural tells: tricolons in series ("X, Y, and Z"), summative closers ("In conclusion…"), colon-titles ("The Future of X: A New Paradigm"), bullet-list overuse, hedged claims without source
- Punctuation tells: em-dash overuse (single signal — not proof; see Ann Handley's published rebuttal); ellipsis outside quotation
- Formula constructions: "It's not just X, it's Y" · "Picture this:" · "Imagine a world where" · "What if I told you" · "Whether you're a seasoned X or a curious newcomer" · "In the realm of" · "Navigating the landscape of"
Diagnose LLM drift quantitatively:
1. grep -c -iE 'delve|leverage|crucial|robust|underscore' across the corpus — frequency ≥1 per 500 words is a strong tell 2. Sentence-length σ < 4 across a 100-sentence window — uniformity is a stronger tell than any single lexical signal 3. n-gram comparison between the brand's pre-AI corpus and post-AI corpus — divergence in top trigrams flags drift
Detection is unreliable as a single source of truth. Use these as triage, not verdict. The Stanford HAI / Liang et al. (2023) work showed GPT detectors misclassify TOEFL essays by non-native English writers at headline rates above 60%. Treat any single signal as suspicion, not proof.
Phase 8 — Render PROSE.md
Use the hybrid template in references/prose-md-template.md:
1. Narrative sections for each layer + policy (the _why_ and the _how_) 2. Do/don't tables as an annex (the quick-reference scan layer) 3. Sample bank: ≥10 before/after pairs, ≥3 exemplar pieces if provided, hook bank and closing bank cross-referenced from samber/cc-skills@copywriting-hooks / @copywriting-cta 4. Cross-references to TONE.md and SOUL.md (read together, not in isolation) 5. Versioning footer: semver, date, owner, changelog stub
---
ADAPT workflow
Take an existing PROSE.md and project it onto a new channel grouping.
1. Read the existing PROSE.md. 2. Ask the user: target channel grouping (long-form / social / email / marketing copy), and optionally a specific platform within the grouping for tighter overrides. 3. Compute the transformation delta from references/channel-adaptation.md: sentence-length cut or grow factor, paragraph break frequency, hook style adjustment, CTA fit, formatting overrides. 4. Emit a CHANNEL OVERRIDE — <grouping> section appended to PROSE.md, or a standalone PROSE-<grouping>.md if the user prefers a separate artifact. Why offer both: content teams that publish across many channels prefer one master file; ghostwriting agencies handling a single channel prefer per-channel files. 5. Cross-reference back to the original PROSE.md for fields unchanged.
---
AUDIT workflow
Extract current prose patterns from a corpus before codifying. Empirical patterns beat invented ones.
1. Take the corpus (folder of .md / .txt or list of URLs). 2. For corpora > 50 pieces, parallelize: spin up to 5 sub-agents via the Agent tool, splitting the corpus by date range, channel, or author. Each agent reports back with the same metrics. Why parallel: sequential reading on a 200-piece corpus is slow and runs out of context; parallel sub-agents read independently and synthesize. 3. Compute (per references/audit-tools.md):
- Mean sentence length and distribution
- Top 50 lexemes, top bigrams and trigrams
- Banned-word and AI-tell frequency
- Em-dash count per 1,000 words
- Opening pattern map (first 50 words of 30 pieces, side by side)
- Closing pattern map
4. Run an adversarial reading pass on 3–5 representative pieces — challenge the assumption that they work. Mark every sentence that doesn't earn its place, every unanswered reader question, every moment authority collapses, every paragraph where a reader would disengage. See references/audit-tools.md for the methodology. 5. Sort findings into four buckets: signature (recurring, distinctive, working) · default (recurring, generic, neutral) · noise (inconsistent, accidental, weak) · liability (recurring, actively harming credibility or engagement — the adversarial pass surfaces these). 6. Produce AUDIT-MEMO.md (5–10 pages: quantitative tables + qualitative annotated samples + "keep, kill, differentiate" summary). Feed into BUILD Phase 3.
---
Output format
PROSE.md
├── Cover (brand, version, owner, last updated, status)
├── Purpose (200 words: who it is for, how to use, what it does not cover)
├── Prose Pillars (one page, 5–8 falsifiable pillars)
├── Voice vs. Tone note (one paragraph)
├── 1. Lexicon (narrative + do/don't annex)
├── 2. Syntax
├── 3. Rhythm
├── 4. Structure
├── 5. Voice Markers
├── 6. Punctuation Policy
├── 7. Formatting Policy
├── 8. Channel Overrides (one section per in-scope grouping)
├── 9. Cultural & Linguistic Adaptation
├── 10. Anti-LLM Countermeasures
├── 11. Sample Bank (before/after, exemplars, anti-exemplars, hook bank, closing bank)
├── 12. Ghostwriting Addendum (per principal — optional)
├── Annex A: Do/Don't quick reference (all layers, scannable)
└── ChangelogA complete PROSE.md is 20–60 pages depending on category coverage and channel scope. Resist the urge to maximize length — Siemens reduced their brand guidelines from 2,750 to 250 pages because enforceable density beats exhaustiveness. Aim for the density that an editor can apply line by line; cut anything an editor cannot turn into a concrete edit.
---
Reference files (load on demand)
| File | When to read |
|---|---|
| discovery-questions.md | During Phase 1 interview |
| five-layers.md | During Phase 3 codification |
| category-playbooks.md | During Phase 2 after category detection |
| channel-adaptation.md | During Phase 5 and all ADAPT invocations |
| anti-patterns.md | During Phase 7 and AUDIT mode |
| multilingual.md | During Phase 6 when brand operates in EN/FR |
| prose-md-template.md | During Phase 8 render |
| brand-atlas.md | During Phase 2 archetype matching |
| audit-tools.md | During AUDIT mode and Phase 3 corpus diagnosis |
---
Disclaimer
This skill is not exhaustive. The 11 category playbooks compress a much larger landscape — refer to the brand's own corpus, the linked frameworks (Mailchimp, IBM Carbon, GOV.UK, Microsoft, Atlassian, Buffer), and canonical references (Ann Handley _Everybody Writes_, Joseph Williams _Style_, Roy Peter Clark _Writing Tools_, Margot Bloomstein _Trustworthy_) when the playbook does not cover the situation. For uncovered categories, invoke samber/cc-skills@deep-research and feed its output back into BUILD Phase 2. Prose guides decay; a PROSE.md not re-audited every 12 months is a snapshot, not a living document.
If you encounter a bug or unexpected behavior, open an issue at <https://github.com/samber/cc-skills/issues>.
Anti-Patterns and AI Tells
The dominant prose-drift risk in content factories is convergence on LLM-default register. This file inventories the patterns to ban explicitly in PROSE.md and to check for in AUDIT mode.
Refresh yearly. Lexical and structural tells evolve as models change.
Lexical tells
Words and phrases that appear at elevated frequency in LLM-generated text. Roger J. Kreuz (_The Conversation_, 2025) documented "a dramatic increase in relatively uncommon words, such as 'delves' or 'crucial,' in articles published in scientific journals over the past couple of years."
Single-word tells
| Word | Common in LLM output as | Replacement strategy |
|---|---|---|
| delve | "let's delve into…" | "examine", or just say the thing |
| leverage | "leverage X to achieve Y" | "use", or restructure to active verb |
| crucial | "this is crucial" | drop or be specific about why |
| robust | "robust framework / system" | name the property concretely |
| underscore | "this underscores the importance of…" | "shows" / "proves" / "means" |
| navigate | "navigate the complexities of…" | drop the navigation metaphor; say the thing |
| seamlessly | "integrates seamlessly" | replace with what it actually does |
| vibrant | "vibrant community" | give a concrete signal of vibrancy |
| dynamic | "dynamic landscape" | be specific about what changes |
| embark | "embark on a journey" | drop the journey metaphor |
| foster | "foster collaboration" | "build" / "support" / "encourage" |
| harness | "harness the power of" | drop the harnessing metaphor |
| myriad | "a myriad of options" | "many" / specific number |
| tapestry | "rich tapestry of" | drop the tapestry metaphor |
| paradigm | "new paradigm" | "model" / "approach" — usually drop |
| pivotal | "pivotal moment" | be specific about why it pivoted what |
| holistic | "holistic approach" | name the components |
| synergize | n/a | banned in any register |
| utilize | "utilize" | "use" |
| facilitate | "facilitate" | "help" / "let" / "enable" |
| commence | "commence operations" | "start" / "begin" |
| furthermore | sentence opener | replace with "and" / drop / restructure |
| moreover | sentence opener | replace with "also" / drop |
| notwithstanding | "notwithstanding the…" | "despite" |
Phrase tells
- "It's not just X, it's Y" — the formula construction; LLMs emit this 5–10× human baseline rate
- "In conclusion,…" — summative closer; humans usually skip it
- "It's important to note that…" — hedging that adds nothing
- "It's worth noting that…" — same
- "Crucially,", "Notably,", "Importantly,", "Essentially," — sentence-opener hedges
- "Picture this:" / "Imagine a world where" / "What if I told you" — manufactured-curiosity openings
- "Whether you're a seasoned X or a curious newcomer" — audience-segmenting filler
- "In the realm of" / "Navigating the landscape of" — empty scene-setting
- "Unlock the power of" / "Dive into" / "Buckle up" / "Let's dive in" — hype openers
- "From the comfort of your own home" — copywriting cliché
- "At the end of the day" — empty connector
- "Last but not least" — banned in lists
- "In today's fast-paced world" / "In an era of" — universally banned opener
- "Hope this helps!" — chatbot signature
French phrase tells
- "Dans un monde en constante évolution" — universally banned French opener
- "Plongez dans…" — equivalent of "dive into"
- "Découvrez comment…" — over-used in French AI output
- "Par ailleurs,…" / "Notamment,…" — over-used sentence openers
- "Il est crucial de…" — French equivalent of "it is crucial to"
- "À l'heure du tout-numérique" / "À l'ère de l'IA" — universally banned openers
- "N'hésitez pas à…" — chatbot-flavored closing
Structural tells
Isabel Al-Dhahir (_Verdict_, 2024): "AI models also overuse tricolons, a rhetorical device that consists of a series of three parallel words, phrases, or clauses."
| Pattern | Why it's a tell | Counter |
|---|---|---|
| Tricolons in series ("X, Y, and Z") | LLMs default to threes everywhere; humans vary list lengths | Limit lists of 3 to 1 per 1000 words; use 2-item or 4-item lists elsewhere |
| Colon-titles ("The Future of X: A New Paradigm") | Over-represented in LLM headlines | Use single-clause titles; drop the colon |
| Summative closers ("In conclusion,…") | Humans skip this in 80%+ of pieces | End on the thing itself; no recap |
| Bullet-list overuse | LLMs convert prose to bullets reflexively | Audit: bullet-density per piece. Cap. |
| Hedged claims without source | "It is widely believed that…" / "Many experts agree…" | Force citations or remove |
| Symmetric paragraph lengths | Uniform 4-sentence paragraphs across the piece | Force variance via the rhythm rules |
| Anaphora overuse | Repeated openers in clusters | Reserve anaphora for closings only |
| "Not only X, but also Y" | LLM-favored coordination | Use "X. And Y." or restructure |
Punctuation tells
The em-dash debate is contested. Ann Handley's published rebuttal (_The Em Dash Is NOT an AI Tell_, annhandley.com) argues it is unreliable as a single signal. Treat as suspicion, not proof.
| Punctuation pattern | Strength of signal |
|---|---|
| Em-dash density > 1 per 200 words | Weak — humans vary widely |
| Ellipsis outside direct quotation | Moderate |
| Parenthetical asides > 1 per 200 words | Weak |
| Em dash + tricolon + summative closer in same piece | Strong |
Statistical tells
GPTZero (Edward Tian) frames detection as perplexity (how predictable to a language model) plus burstiness (sentence-length variance).
- Low perplexity correlates with AI generation. Hard to measure without tooling.
- Low burstiness = uniform sentence length. Measurable: σ < 4 words per 100-sentence window suggests automation. σ ≥ 6 is the human baseline target.
Detection unreliability
Liang et al. (2023), _GPT detectors are biased against non-native English writers_ (_Patterns_, Cell Press; Stanford HAI): "These platforms incorrectly labeled more than half of the essays as AI-generated, with one detector flagging nearly 98%" of TOEFL essays by non-native English writers.
Operational consequence: use detectors as triage, not verdict. Codify rules LLMs cannot follow by default — that is the durable defense, not a detector arms race.
Counter-measures for content factories
1. Codify rules that LLMs cannot follow by default: idiosyncratic lexicon, named taboos, sentence-length variance targets, specific voice markers. 2. Use the PROSE.md as the AI system prompt; supplement with the brand's published corpus as few-shot examples. 3. Editor reviews 100% of AI-assisted content with the audit scorecard. 4. Run regular n-gram comparison between the brand's pre-AI corpus and post-AI corpus to detect lexical drift. 5. Invoke samber/cc-skills@humaniseur-fr (for French) or an equivalent humanizer skill as a final pass on AI-assisted drafts. Note: humanizers do not replace the prose guide — they scrub the patterns the guide already banned.
Cliché opener inventory
A list to embed in the brand's taboos section. These all immediately disqualify the writer:
English:
- "In today's fast-paced world…"
- "Have you ever wondered…?"
- "Did you know…?"
- "What if I told you…?"
- Dictionary opener played straight ("Productivity, defined as…")
- "In this article, I'll discuss…"
- "I'm not an expert, but…"
- Three rhetorical questions in a row
- "Imagine waking up…" without a specific scene
- "Hot take:", "Unpopular opinion:"
- "At [Company], we believe…"
- "Recently,…" without a specific date
- "You're not alone."
- "We've all been there."
- "Buckle up,"
- Misattributed Einstein / Seneca / Confucius / Bouddha quotes
French:
- "À l'heure du tout-numérique"
- "À l'ère de l'IA"
- "Dans un monde où…"
- "Vous êtes-vous déjà demandé…?"
- "Dans cet article, nous allons voir…"
- "Je ne suis pas spécialiste mais…"
- "Voici la vérité que personne ne veut entendre…"
- "Récemment,…" sans date précise
- "Chez [Entreprise], nous pensons…"
- "Cher lecteur," (in newsletter)
- "Les études montrent que…" sans source
- "90% des gens…" sans source
Diagnose
When auditing a piece for AI tells:
1. grep -c -iE 'delve|leverage|crucial|robust|underscore|seamlessly|navigate|harness|foster|embark|myriad|tapestry|holistic|paradigm|utilize|facilitate|commence' — count single-word tells. > 1 per 500 words = strong tell. 2. Count em dashes per 1000 words. > 5 = check the context. 3. Compute sentence-length σ on a 100-sentence sample. σ < 4 = robotic cadence. 4. Search for the formula constructions ("It's not just X, it's Y"; "Whether you're a seasoned X"). Any hit = rewrite. 5. Check tricolon density per 500 words. > 3 = AI-favored.
Audit Tools
For AUDIT mode and for Phase 3 corpus diagnosis. The goal: compute objective prose metrics on a corpus so that the prose guide's targets are empirical, not invented.
These tools do not produce prose recommendations directly. They surface signals; the model interprets them and codifies the rules.
Readability
Hemingway Editor
Browser tool at hemingwayapp.com (or the desktop / pasteable web version). Highlights:
- Sentences hard to read (yellow) and very hard (red)
- Passive voice
- Adverbs
- Complex phrases ("utilize" → "use")
- Reading grade level
Usage: paste 1000 words. Read off the grade level. Targets per category:
| Category | Target grade |
|---|---|
| B2C / consumer brand | 6–9 |
| B2B SaaS | 9–12 |
| NGO / nonprofit | 7–10 |
| Industry / deep-tech | 12–16 |
| Consulting | 11–14 |
Why grade level matters: it correlates with sentence length, syllable count, and Latinate ratio. A grade level mismatched with audience expertise predicts drop-off.
Plain Language Commission
For UK English specifically. Maps to the same metrics with slightly different targets.
Banned-word enforcement
Vale
Vale (vale.sh) is the dominant prose linter. It applies YAML rules to Markdown / plain text and flags violations.
Sample Vale rule (banned word):
extends: existence
message: "Banned word: '%s'. Use a plain alternative."
level: error
tokens:
- delve
- leverage
- crucial
- robust
- underscore
- seamlessly
- navigate
- harness
- foster
- embark
- myriad
- tapestry
- paradigm
- utilize
- facilitate
- commencePlace under .vale/styles/Brand/BannedWords.yml. The brand can extend with category-specific bans (e.g., "synergize" for consulting).
Why Vale, not LanguageTool, for banned words: Vale is built for prose-rule enforcement; LanguageTool is built for grammar correction. Different tools, different jobs.
LanguageTool
LanguageTool (languagetool.org or the open-source self-host) catches grammar errors, typos, and weak constructions. Strong on:
- Subject-verb agreement
- Passive voice (alerts; does not auto-rewrite)
- Wordiness suggestions
- Repeated words
Configuration: import the brand's banned-word list as a custom dictionary; whitelist intentional brand terms (the lowercase Innocent, the all-caps Apple SHORTCUTS, etc.).
Grammarly Business
Same job class as LanguageTool, commercial. Stronger UI for distributed writers. Custom dictionary supports the brand's banned and required terms. Browser plugin enforces in-flow.
Quantitative metrics
Mean sentence length and distribution
A short Python snippet for a quick audit:
import sys
import re
from statistics import mean, stdev
text = sys.stdin.read()
# Split on sentence terminators, ignoring abbreviations
sentences = re.split(r'(?<=[.!?])\s+', text)
sentences = [s.strip() for s in sentences if s.strip()]
lengths = [len(s.split()) for s in sentences]
print(f"Sentences: {len(sentences)}")
print(f"Mean length: {mean(lengths):.1f} words")
print(f"Std dev: {stdev(lengths):.1f}")
print(f"Min / max: {min(lengths)} / {max(lengths)}")
print(f"≥ 25 words: {sum(1 for l in lengths if l >= 25)} ({100*sum(1 for l in lengths if l >= 25)/len(lengths):.1f}%)")
print(f"≤ 8 words: {sum(1 for l in lengths if l <= 8)} ({100*sum(1 for l in lengths if l <= 8)/len(lengths):.1f}%)")Run on each piece in the corpus. Aggregate per channel.
Type-token ratio
Vocabulary diversity. Lower TTR = more repetition (often a sign of LLM-flat lexicon). Compute as unique_words / total_words on a 1000-word window. Healthy human B2B prose lands at 0.40–0.55; LLM-default sits at 0.30–0.42 due to favored vocabulary clustering.
n-gram frequency
For lexical drift detection. Compare two corpora (e.g., pre-AI vs post-AI):
from collections import Counter
def ngrams(text, n=3):
tokens = text.lower().split()
return Counter(' '.join(tokens[i:i+n]) for i in range(len(tokens) - n + 1))
pre = ngrams(open('pre_ai_corpus.txt').read())
post = ngrams(open('post_ai_corpus.txt').read())
# Trigrams that surged
diff = [(t, post[t] - pre[t]) for t in post]
diff.sort(key=lambda x: -x[1])
for t, d in diff[:30]:
print(f"{d:+5d} {t}")Top surging trigrams flag the brand's drift vectors.
Banned-word frequency
# Count banned-word hits across a corpus
grep -roEic 'delve|leverage|crucial|robust|underscore|seamlessly|navigate|harness|foster|embark|myriad|tapestry|holistic|paradigm|utilize|facilitate|commence' ./corpus/ | sort -t: -k2 -nrHits per 500 words is the relevant rate.
Reading-aloud audit
The lowest-tech tool, often the most revealing. Read the piece aloud:
1. Mark every breath point. 2. Mark every sentence where you stumble or have to re-read. 3. Mark every monotonous stretch (more than 4 consecutive medium sentences). 4. Mark every accidental rhyme or alliteration.
A trained editor catches in 5 minutes what statistical tools miss. Codify the read-aloud audit as part of the per-piece QA checklist for high-stakes content (pillars, executive op-eds, keynotes).
n-gram comparison for ghostwriting
For ghostwriting voice-matching audits: compute n-gram overlap between the principal's authentic writing/speech corpus and the ghostwriter's drafts. Low overlap on signature n-grams = the ghostwriter has flattened the principal's idiolect.
This is the operational test for the "thinking translation problem" — when the ghostwriter reproduces the client's sentence structure and vocabulary but misses the actual operational insight only the client could have.
Web-based readability tests
For a quick sanity check without local tooling:
- Hemingway Editor (hemingwayapp.com) — grade level, complex sentences, passive voice, adverbs
- Datayze Sentence Length Checker (datayze.com/sentence-length-checker) — distribution histogram
- WebFX Readability Test (webfx.com/tools/read-able) — multiple readability scores
Vale + CI integration
For content factories at scale, wire Vale into the editorial CI:
# .vale.ini
StylesPath = .vale/styles
MinAlertLevel = warning
[*.md]
BasedOnStyles = Brand, Microsoft, write-goodRun on every pull request to the content repo. Fail the build on Vale errors. Why: the prose guide is enforced by tooling, not by editor fatigue. Editors should review judgment calls, not catch banned words.
Adversarial reading
Quantitative tools surface signals. Adversarial reading surfaces what the numbers miss: the sentences that don't earn their place, the moments authority collapses, the reader questions that go unanswered.
Core posture: the writer already believes the draft works. Challenge that assumption. Read to find what fails, not to confirm what succeeds.
Protocol (per piece)
Read the piece once without stopping. Then re-read and mark:
1. Dead weight — sentences or phrases that could be deleted without information loss. Count them. A ratio above 15% signals a draft, not a final piece. 2. Authority collapse — claims that invite "says who?", statistics without sources, analogies that don't hold, jargon used to signal expertise rather than convey meaning. 3. Reader dropout points — paragraphs where a reader would disengage: slow accumulation with no payoff, five consecutive medium-length sentences with no breath point, transitions that require re-reading. 4. Unanswered questions — the "so what?" question every factual claim generates. If the paragraph raises a question and the next paragraph doesn't resolve it, the structure is broken. 5. Distance incoherence — sudden shifts in psychic distance (see five-layers.md § 4.11) with no structural reason (e.g., a close-second-person hook that snaps to third-person-corporate in paragraph two).
Critique dimensions for brand prose
Adapted from fiction critique methodology (haowjy/creative-writing-skills@prose-critique):
| Dimension | For brand prose | Key question |
|---|---|---|
| Structure | Piece-level coherence, pacing, payoff mechanics | Does each section earn its length? Does the opening pay off at the close? |
| Voice | POV stability, dialogue effectiveness, implicit meaning | Does the brand voice hold across the full piece, or drift mid-article? |
| Prose | Sentence-level craft, rhythm, word repetition, descriptor-to-narrative ratio | Are there more than 3 consecutive sentences of the same type? |
| Brand persona | Motivational consistency of the brand character | Does the brand's stated archetype match the brand's actual prose behavior? |
| Continuity | Factual accuracy, claim consistency within the piece | Does the piece contradict itself across sections? |
Sorting findings
Map each finding to a bucket:
- Signature — recurring and working; codify as rules to preserve
- Default — recurring and neutral; decide whether to keep or differentiate
- Noise — accidental and inconsistent; no action needed
- Liability — recurring and harmful to credibility or engagement; codify as explicit prohibitions
The liability bucket is what adversarial reading surfaces that metrics miss.
Limits of automation
These tools surface signals; none replaces editorial judgment. A piece can pass every quantitative check and still read flat — because the voice markers are missing, or the structure is wrong, or the hook doesn't earn the close. Use audit tools as the first filter, then ship to human editorial review for the rest.
Brand Atlas
A short, opinionated catalogue of brands whose prose is identifiable in a blind test. Use during Phase 2 archetype matching: ask "which of these does the brand most resemble?" and "where does it want to differentiate?"
Each entry: one-line summary of the diagnostic feature.
Consumer brands
Mailchimp — short, declarative, sentence-case, contractions, Oxford comma, plain Anglo-Saxon verbs, sparing humor, never patronizing. _Diagnostic_: sounds like a colleague writing an email at 3pm.
Innocent Drinks — ultra-short sentences (frequently 4–8 words), product-as-narrator first person, lowercase brand name, hand-written-feeling asides, fruit puns rationed. _Diagnostic_: a sentence that ends in "Win." or a parenthetical aside in product voice.
Patagonia — long, journalistic, magazine-feature openings; named human protagonists (Chilean fishermen, female carpenters, Japanese sake brewers); data points embedded in narrative; declarative ethical claims. _Diagnostic_: a story about a named person with a place name in the dek.
Liquid Death — maximalist horror-genre register, mock-tabloid microcopy, all-caps headlines, sentence fragments as exclamations, satirical "About" pages. _Diagnostic_: a product description that reads like a B-movie pitch.
Oatly — handwritten-style asides, faux-naïve voice, self-deprecation, parenthetical meta-commentary on its own marketing.
Apple — terminal sentence economy and sentence-case discipline.
B2B / SaaS brands
Stripe — footnoted, narrative-memo prose. Patrick Collison structures emails like research papers with footnotes. _Diagnostic_: a paragraph that pre-empts an objection in parentheses or a footnote. Stripe defaults to writing over slide decks across the entire company.
Linear — terse, declarative, no marketing adjectives, single-sentence value propositions, present-tense product claims, dense product screenshots in lieu of explanation. _Diagnostic_: a sentence with no adjective.
Slack — friendly imperatives, microcopy as small talk, sentence-case headings, contractions, frequent second person. _Diagnostic_: a UI string that begins with a verb and feels like a coworker.
Notion — didactic-but-warm, lists everywhere, second-person, "you can…" constructions, definitions before features. _Diagnostic_: help-doc cadence even in marketing copy.
Basecamp / 37signals — opinionated, contrarian, short paragraphs, frequent one-sentence paragraphs, dialectical structure (claim/counterclaim/resolution). _Diagnostic_: a one-sentence paragraph that contains an opinion.
Mailerlite — clear, didactic, list-heavy, second-person, low jargon, screenshot-rich. _Diagnostic_: a numbered list that is the article.
Social media voices
Wendy's social — punchline-first, sentence fragments, callbacks to previous posts, no hashtags. _Diagnostic_: a reply that is funnier than the prompt.
Duolingo social — chaos register, lowercase, ironic threats, owl-as-narrator. _Diagnostic_: a post that would be a fireable offense at any other brand.
Reference frameworks
The published style guides worth reading in full before codifying:
Nielsen Norman Group, _The Four Dimensions of Tone of Voice_ (Moran, 2016) — funny↔serious, formal↔casual, respectful↔irreverent, enthusiastic↔matter-of-fact. Use as sanity check; not a prose rule.
Mailchimp Content Style Guide (styleguide.mailchimp.com) — most-copied open template. Take the structure (Voice and tone, Grammar and mechanics, Writing about people, Writing for accessibility).
GOV.UK Style Guide — the discipline of plain language. Targets reading age 9 across the site; maintains a strict A–Z banned-word list. Take the discipline of a maintained A–Z list.
Microsoft Writing Style Guide (learn.microsoft.com/style-guide) — sentence case in headings, Oxford comma, contractions allowed, "Write short, simple sentences." Mean target 15–20 words per sentence.
IBM Carbon Design System (carbondesignsystem.com/guidelines/content/overview) — codifies voice as a nine-attribute checklist. Take the rare technique of writing voice as a checklist an editor can apply line by line.
Atlassian Design System (atlassian.design/content) — "be bold, be optimistic, be practical, with a wink" tetrad, paired with the internal "Voicify" sliding-scale tool. Take the situational-flex model for newsletters and product-led content.
Buffer Style Guide (buffer.com/resources/style-guide) — voice attributes "relatable, approachable, genuine, inclusive". Product copy rules unusually prescriptive: "Invite the customer to take an action. Never command." Take the prescriptive product-copy block.
Canonical writing books
For the brand owner and editor to read once, not the writers:
- Ann Handley, _Everybody Writes_ (2nd ed., 2022) — the 17-step "Writing GPS", the "ugly first draft" doctrine, "Create reading momentum."
- William Zinsser, _On Writing Well_ — the pruning doctrine ("clutter is the disease of American writing"). Operationalize as a 20% cut rule on every draft.
- Strunk & White, _The Elements of Style_ — rule 17 ("Omit needless words"), rule 11 ("Use the active voice"). Overrideable defaults.
- Joseph Williams, _Style: Lessons in Clarity and Grace_ — the character-action principle (subjects should be characters, verbs should be actions). The most useful syntactic rule for technical writers escaping nominalizations.
- Roy Peter Clark, _Writing Tools_ — Tool 2 ("Order words for emphasis": strongest word at the end, second strongest at the start) and the parallelism toolkit.
- Heath brothers, _Made to Stick_ — the SUCCES heuristic for openers and closers, especially Concreteness.
- Margot Bloomstein, _Content Strategy at Work_ and _Trustworthy_ — the BrandSort message-architecture method, the prerequisite to a prose guide.
- Nicole Fenton and Kate Kiefer Lee, _Nicely Said_ — the simple style-guide template at the end of the book is the minimum viable prose guide.
Corpus linguistics methods
For quantitative audits, the relevant methods:
- Type-token ratio — vocabulary diversity per piece
- Mean sentence length distribution — central rhythm signal
- n-gram frequency — top bigrams and trigrams; drift detection
- POS-tag profiles — verb/noun/adjective ratios; detects nominalization drift
- Banned-word frequency — direct enforcement metric
These are the only way to detect prose drift in a large content operation. See audit-tools.md for the actual tools.
How to use this atlas in Phase 2
1. Ask the brand owner which 1–3 brands they admire most (not necessarily in their category). 2. Ask which 1–3 brands they actively want to avoid sounding like. 3. Map both sets to the entries above where possible; for unrecognized brands, run an inline corpus skim. 4. Codify the "differentiate" axis: which of the loved brand's signature moves to borrow; which of the avoided brand's defaults to ban.
The atlas is not exhaustive. A brand that does not resemble any entry here is interesting — codify what they actually do, then file it under a new entry for next time.
Category Playbooks
Eleven playbooks. Each one compresses corpus evidence for the category — load the relevant playbook in Phase 2, apply its defaults in Phase 3.
Format per playbook: Optimize for · Prose characteristics · Anti-patterns · Lexicon · Syntax · Rhythm · Structure · Reference brands.
For categories not in this file, invoke samber/cc-skills@deep-research to map the category's prose conventions before codifying.
1. B2B (SaaS, enterprise, tech)
- Optimize for: clarity + credibility
- Prose characteristics: declarative-default; concrete examples within 100 words of any abstract claim; numbers, dates, customer names; precise terminology; explicit connectors (because, therefore, in contrast); structured arguments
- Anti-patterns: vague benefit-speak ("transform your business"); adjective stacking ("powerful, scalable, intelligent"); long sentences with passive subjects ("It is believed that…"); thought leadership without an actual thought
- Lexicon: technical terms permitted at a known specificity (latency, throughput, TCO, MRR) with glosses for non-expert readers; ban marketing intensifiers (best-in-class, world-class, cutting-edge); use product names exactly as registered
- Syntax: mean 14–18 words; allow longer for technical claims with multiple qualifications; subject-verb proximity (no buried verbs); active default, passive permitted in technical impersonal claims
- Rhythm: alternate short claim with longer evidence; one breath sentence (≤ 8 words) every 4–6 sentences; H2/H3 every 200–300 words
- Structure: TL;DR or summary up top; inverted pyramid; numbered lists for procedures; tables for comparisons; explicit headings ("How it works", "Why it matters", "What to do")
- Reference brands: Stripe (footnoted memos), Linear (terse declaratives), Notion (didactic-warm), Atlassian (clear product prose)
- France-specific: French B2B SaaS audiences in English expect lower tolerance for hyperbole and higher tolerance for technical depth than US norms
2. B2C (consumer products)
- Optimize for: emotional clarity + memorability
- Prose characteristics: shorter sentences; second person; concrete sensory verbs (taste, feel, see); product as protagonist; benefit-first leads; one idea per paragraph
- Anti-patterns: B2B jargon (solutions, platforms); features without sensory translation; long sentences; passive constructions; abstract nominalizations
- Lexicon: Anglo-Saxon verbs over Latinate; no acronyms without explanation; product names always capitalized as registered
- Syntax: mean 10–14 words; sentence fragments permitted in body for emphasis; imperatives common in CTAs
- Rhythm: punchy openings; one-sentence paragraphs allowed; lists kept to 3 items (Rule of Three)
- Structure: lead with the feeling, then the feature; testimonials and reviews integrated; product photography as part of the structural rhythm
- Reference brands: Innocent Drinks, Mailchimp consumer copy, Apple consumer pages
3. Consumer brand (lifestyle, DTC)
- Optimize for: distinctiveness + memorability
- Prose characteristics: high voice-marker density; idiosyncratic signature moves; willingness to break grammar rules deliberately (Innocent's lowercase i, Oatly's run-on asides); brand-as-character voice; meta-commentary on its own marketing acceptable
- Anti-patterns: imitating Innocent without earning it (Nick Asbury's "Wackywriting" critique); cuteness without substance; ironic distance that masks lack of conviction
- Lexicon: brand-owned coinages encouraged (Liquid Death's "Murder Your Thirst", Oatly's "Wow no cow!"); strong opinions encoded in vocabulary
- Syntax: deliberately varied; fragments common; one-sentence paragraphs frequent; rule-breaking permitted but rationed
- Rhythm: highly variable cadence; surprise sentence lengths; punchlines at the end
- Structure: micro-content first (packaging, social); pillars often manifestos rather than how-to
- Reference brands: Innocent Drinks, Liquid Death, Oatly, Patagonia (purpose-led variant), Cards Against Humanity (irreverent variant)
- Caveat: this register collapses without a real product point of view. Liquid Death works because the product (canned water in beer-style cans) is itself a thesis
4. Non-corporate / NGO / non-profit
- Optimize for: dignity + clarity
- Prose characteristics: humans named with consent, not anonymous "beneficiaries"; data embedded in narrative, not opposed to it; specific places and dates; first-person testimonials with full attribution where possible; ethical care in describing vulnerable populations
- Anti-patterns: poverty porn / suffering aesthetics; abstract nouns of suffering ("hunger", "injustice") without grounding; corporate-speak imported from for-profit ("stakeholders", "engagement", "ROI of empathy"); patronizing framings ("giving voice to" instead of platforming)
- Lexicon: people-first language ("people experiencing homelessness", not "the homeless"); strengths-based vocabulary (Charity: Water — focus on hope, not guilt); avoid Latinate abstractions in favor of Anglo-Saxon verbs; maintain a banned list for stigmatizing terms
- Syntax: medium sentence length (15–20 words) to accommodate context; active voice default with named agents; passive when protecting privacy
- Rhythm: narrative cadence; alternating personal scene and aggregate data; explicit "from one person to many" structure
- Structure: scene → person → context → systemic claim → call to action; named impact ("37 wells in 12 villages, serving 4,800 people"); credit lines for community partners
- Reference brands: Charity: Water (hope-not-guilt), Médecins Sans Frontières (clinical-but-human), Oxfam (campaigning), Patagonia (purpose-led)
- France-specific: French NGO register is more abstract and lexically Latinate than English equivalents; English-language adaptations should cut 20% of abstract vocabulary
5. Consulting / professional services
- Optimize for: authority + accessibility
- Prose characteristics: structured arguments with named frameworks; data with citations; institutional "we have observed" or partner-level "I have observed"; concrete client situations (anonymized where required); explicit caveats and conditions
- Anti-patterns: McKinsey-pastiche ("In our experience, leading organizations…") without specifics; consulting bingo (synergies, optimization, transformation); claims without evidence; pseudo-data ("studies show")
- Lexicon: declared framework terms (capitalize when proper noun, lowercase otherwise); banned consultancy jargon list (synergize, operationalize, leverage as verb); precise hedges ("we observed in 7 of 12 engagements", not "often")
- Syntax: longer sentences acceptable (18–22 words mean) given expert audience; subordination acceptable; parenthetical conditions common
- Rhythm: slower cadence than B2B SaaS; long paragraphs of argument (5–8 sentences) acceptable in pillars; broken by data callouts or numbered findings
- Structure: executive summary mandatory; situation–complication–resolution (Minto Pyramid) common; numbered findings; explicit "implications for X" sections
- Reference brands: McKinsey Quarterly (institutional), Bain Insights (data-driven), Stripe Press (technical-cultural), BCG Henderson Institute (academic-adjacent)
6. Product-led (makers, indie hackers, dev tools)
- Optimize for: utility + product as evidence
- Prose characteristics: present-tense product claims; screenshots and code blocks as part of the prose; "show, don't tell" enforced (Linear's approach); CTAs are product demos, not contact forms; documentation tone leaking into marketing copy (positively)
- Anti-patterns: enterprise marketing imported into product-led contexts ("revolutionize your workflow"); long preambles before showing the product; benefits without screenshots
- Lexicon: feature names treated as proper nouns; verbs that map to UI actions (click, select, configure); no industry buzzwords unless category requires
- Syntax: very short, mean 10–14 words; imperatives in how-to content; declaratives for product claims
- Rhythm: claim, screenshot, claim, screenshot; very high signal-to-noise
- Structure: above-the-fold product image plus one-sentence claim; problem-solution-proof below; final CTA = try the product
- Reference brands: Linear, Notion, Vercel, Plausible, Raycast
7. Industry (manufacturing, B2B industrial, deep-tech)
- Optimize for: precision + technical credibility
- Prose characteristics: specifications cited exactly; standards referenced (ISO, IEC, ASTM); named engineers and facilities; metric units; long-form technical depth in pillars with executive-summary openings
- Anti-patterns: consumer marketing language (amazing, incredible); soft claims without spec (high performance); breathy purpose-marketing pasted onto industrial products
- Lexicon: domain-specific terminology required, not avoided; standard abbreviations as norms (kW, MPa, IP67); banned consumer intensifiers; full part numbers and model designations
- Syntax: long sentences tolerated (up to 25-word means in technical sections); passive voice acceptable in scientific impersonal register; precise qualifications ("at 25°C and 1 atm")
- Rhythm: slower; longer paragraphs; data tables interrupt prose
- Structure: GE Reports model is the journalistic benchmark — named protagonist, named place, named challenge, measurable outcome; brand mentioned sparingly
- Reference brands: GE Reports (journalistic storytelling), Siemens (industrial minimalism), Anthropic and OpenAI research blogs (deep-tech variant — long-form, hedged, mission-framed)
- Deep-tech variant: research-blog prose hedges epistemically ("preliminary results suggest"); cites primary sources; reserves superlatives; embeds equations and diagrams
8. Volunteering / community / association
- Optimize for: belonging + actionability
- Prose characteristics: inclusive pronouns (we, our community); concrete next actions per piece (a meetup, a contribution, a vote); volunteer-authored voices given billing; gratitude as a structural element, not closing platitude
- Anti-patterns: corporate-charity register imported from NGOs; insider jargon excluding newcomers; assumed knowledge ("as you know from last month's meeting"); event posts missing basics (date, place, time, who can come)
- Lexicon: declared insider terms with one-line glosses for newcomers; banned gatekeeping vocabulary; inclusive terms (newcomers vs "noobs")
- Syntax: short, conversational; second person plural often appropriate; imperatives for calls to action ("RSVP by Friday")
- Rhythm: brisk; bullet lists for logistics; narrative for stories from members
- Structure: who-what-when-where-why in the first 100 words for events; member spotlights as recurring format; "how to get involved" as standard closing
- Reference brands: open-source community newsletters (Rust Foundation, Kubernetes); volunteer-run conferences (PyCon, GopherCon); local-chapter newsletters
9. Personal branding (individual creators, founders, executives)
- Optimize for: authentic voice fingerprint + cumulative point of view
- Prose characteristics: first person used deliberately; signature openings; short paragraphs on social (1–3 sentences); idiolect preserved (specific filler words, recurring connectors, characteristic metaphors); opinions clearly stated
- Anti-patterns: ghostwritten prose that smooths the principal's idiolect into generic LinkedIn cadence; thought leadership without thoughts; performative vulnerability; AI-template structure (hook + three bullets + question CTA)
- Lexicon: principal-specific. Maintain the principal's actual word list (favorite verbs, characteristic adjectives, terms they refuse to use)
- Syntax: principal-specific. Some founders use long sentences (Paul Graham); others use short (Naval Ravikant). Match the principal's mean sentence length within ±2 words
- Rhythm: principal-specific. Capture 60–90 minutes of recorded speech to extract
- Structure: post archetypes (lesson + story, contrarian opinion, framework, reaction to industry event); each archetype has an established cadence per principal
- Diagnostic test: a blind reader should be able to identify the principal from a paragraph
- Ghostwriting addendum (mandatory for per-principal codification):
1. Corpus capture: 60–90 min of recorded speech + all prior writing (emails, posts, op-eds) 2. Idiolect analysis: mean sentence length, top 50 lexemes, characteristic openings, filler phrases (you know, the thing is), preferred connectors, recurring metaphors, frequent stories, banned topics 3. Codify: 10 signature openings · 5 signature closings · 3 recurring stories with retelling cadence · banned topics · preferred connectors · average post length · line-break convention 4. Calibration: 3 test posts; principal reviews; iterate until "this sounds like me" with no edits 5. Quarterly review with principal
10. Politics / advocacy / public figures
- Optimize for: clarity of position + memorability + ethical durability
- Prose characteristics: declarative-default; concrete promises and concrete constraints; rhetorical devices (anaphora, tricolons, contrast) used deliberately but rationed (overuse signals AI or amateurism); explicit acknowledgement of opposing views before refutation; consistent message across speeches, posts, op-eds
- Anti-patterns: hedging that obscures position; insider procedural language; attack-only register; recycled boilerplate from prior campaigns; superlatives ("most important election of our lifetime") that erode credibility through repetition
- Lexicon: campaign-specific declared vocabulary (signature words and phrases that recur); avoidance of opposition's framing language; inclusive pronouns balanced with specific named groups; banned terms list to prevent gaffes
- Syntax: shorter than other categories (mean 12–16 words); parallelism enforced in promises and value statements; tricolons reserved for keynote moments
- Rhythm: oratorical when intended for speech (read-aloud testing mandatory); written rhythm for op-eds; both should match the principal's natural cadence
- Structure: Aristotle's ethos-pathos-logos still operative. For speeches: hook (named constituent), thesis, three reasons, opposition acknowledged, peroration. For op-eds: argument-led, evidence-rich, concrete proposal
- Reference: Gettysburg Address (272 words, three paragraphs, ten sentences) as the canonical short modern model; Amnesty International and Human Rights Watch for advocacy NGO benchmarks
- Ethical guardrail: stick to what can realistically be accomplished. Misleading voters wins applause in the short term but erodes trust in the long run. Invoke
samber/cc-skills@deep-researchfor jurisdiction-specific legal constraints (campaign finance disclosure, defamation thresholds, electoral codes) before codifying - Note: per-principal customization required, similar to personal branding addendum
11. Internal corporate communication
- Optimize for: trust + comprehension + actionability
- Prose characteristics: plain language; named owners and named deadlines; "what changes for you" sections; explicit acknowledgement of uncertainty during change; consistent voice from leadership across channels (all-hands, intranet, email, Slack)
- Anti-patterns: corporate euphemism for layoffs / restructuring ("synergies", "right-sizing", "transitioning colleagues out"); buried bad news (the real news in paragraph 6); jargon-as-power ("strategic alignment workstream"); ChatGPT-flavored "I am pleased to announce" boilerplate
- Lexicon: employee-facing terms (colleagues, teammates) over HR-coded ones (resources, headcount); concrete role names over abstract function names (engineering team over engineering organization); accessibility-first (avoid acronyms unique to one division)
- Syntax: mean 14–18 words; second person ("you", "your team") in change comms; imperatives for action items
- Rhythm: bullet-heavy for change comms; narrative for context and rationale; clear delineation between "what" / "why" / "what changes for you" / "what to do next"
- Structure: lede paragraph carries the news (no throat-clearing); FAQ section for change announcements; named contact for follow-up questions; explicit timeline
- Reference: Patrick Collison's leaked Stripe internal memos style (research-paper structure with footnotes); Basecamp / 37signals all-hands updates (contrarian-but-warm); Slack's internal comms playbook (transparency by default)
- Critical: this category collapses into HR-speak more reliably than any other. The prose guide must explicitly ban the corporate-comms cliché set (cascading communication, leveraging synergies, key stakeholders, going forward, at this time)
Multi-category brands
A brand may sit in two categories simultaneously (e.g., a B2B SaaS that markets like a consumer brand — Notion, Linear). In that case, write one PROSE.md per audience segment, not a blended guide. A blended guide collapses to the lowest common denominator and loses both audiences. Maintain a mapping document for shared pillars and divergent rules.
Disclaimer
These playbooks compress the dominant patterns of each category as observed in the source research. They are starting points, not endpoints. Audit the brand's actual corpus (via AUDIT mode) before locking the category defaults — empirical patterns beat compressed ones.
Channel Adaptation
Transformation rules between the four generic channel groupings. Used in BUILD Phase 5 (to populate channel-overrides sections in PROSE.md) and in ADAPT mode (to project an existing PROSE.md onto a new channel).
Generic groupings keep PROSE.md portable: adding a new platform within a grouping does not require re-codification. Platform-specific quirks (LinkedIn's algorithm, Substack's paywall, X's reply economy) live in the downstream writer skills, not here.
The four groupings
| Grouping | Covers |
|---|---|
| Long-form articles | Blog posts, pillar pages, evergreen essays, technical deep-dives, opinion essays (Substack web post, Medium, dev.to, own blog — same group) |
| Social posts | LinkedIn, X / Twitter, Bluesky, Threads, TikTok captions, Mastodon |
| Email & newsletter | Newsletter issues (Substack email, ConvertKit, Beehiiv), transactional, drip sequences, lifecycle |
| Marketing copy | Landing pages, ad copy, press releases, podcast show notes, video scripts, sales decks |
Channel deltas
Each row is a transformation rule. When ADAPT mode projects a PROSE.md from one grouping onto another, walk this table top-down.
| Dimension | Long-form | Social | Email & newsletter | Marketing copy |
|---|---|---|---|---|
| Mean sentence length | 14–18 words (B2C 10–14; consulting 18–22) | 8–12 words | 10–14 words | 8–12 words |
| Paragraph length | 2–5 sentences, max 80 words | 1–3 sentences, frequent breaks | 1–2 sentences | 1–3 sentences |
| Hook style | Scene · contrarian · stat · concrete detail · question | Bold claim · direct problem · concrete detail · counterintuitive | Personal frame · curiosity gap · open loop | Promise · direct problem · authority · benefit |
| Body structure | TES, PEEL, inverted pyramid | Hook → re-hook → ABT (And/But/Therefore) → CTA | Personal lede → 3–5 paragraphs → P.S. | Above-fold claim → proof → CTA |
| List use | Bulleted lists within prose, max 7 items | Numbered or bulleted, heavy use, max 5 items | Sparse — break with bold instead | Bullet-heavy for features |
| Heading frequency | Every 200–350 words (H2/H3) | None (line breaks instead) | Light — bolded section headers | One H1, optional H2 per section |
| Closing | Practical next step · callback · restated stake | Specific reply prompt · no open-ended questions | Single primary CTA in P.S. or final paragraph | Direct action button + risk reversal |
| Voice marker density | Moderate (1–2 signature moves per piece) | High (1 marker per post — the brand's "tell") | Moderate (personal opener as marker) | Low (consistency over personality) |
| Reading time | 3–15 min | 15–60 sec | 2–5 min | 30–90 sec |
| Tone register | Voice unchanged, tone matter-of-fact to enthusiastic | Voice unchanged, tone casual-confident | Voice unchanged, tone personal-intimate | Voice unchanged, tone direct-confident |
Transformation rules (long-form → other channels)
The most common ADAPT direction. Pillar articles are usually the source of truth; other channels derive.
Long-form → social post (200–300 words; one platform-specific tweet/post)
1. Extract the single most counterintuitive claim from the article. 2. Lead with it as the hook (re-engineer per samber/cc-skills@copywriting-hooks). 3. Cut sentence length to ~60% of the long-form mean. 4. Break paragraphs to 1–3 sentences. 5. Add line breaks for scannability. 6. CTA: specific reply prompt, comment, link to long-form, or no-CTA. Avoid open-ended questions — they reduce action. 7. Strip qualifiers ("often", "in general", "typically") — social rewards confidence.
Long-form → newsletter issue (500–1500 words)
1. Open with the personal or editorial frame ("This week I've been thinking about…", "A reader emailed me…"). 2. Compress argument to 3–5 paragraphs, one idea per paragraph. 3. Include one named data point within the first 200 words. 4. Single primary CTA — in P.S. or final paragraph. Newsletter readers reward restraint. 5. Cut images sparingly — many email clients block them by default. Lead with strong cover image if used. 6. Subject line: declarative, specific. Avoid clickbait — newsletters live on trust.
Long-form → marketing copy (landing page, ad, press release)
1. Strip narrative scaffolding (anecdotes, scene-setting). Marketing copy lives above the fold. 2. Lead with the promise or direct problem. The reader must know within 3 seconds whether to keep reading. 3. Translate features → benefits → outcomes. 4. Add risk reversal (guarantees, social proof, testimonials). 5. Single CTA per page. Multiple CTAs split conversion. 6. Banned in marketing copy unless brand-specific: rhetorical questions, hedges, scene openings.
Long-form → email & newsletter drip sequence
1. Split the article's arguments across N emails — one argument per email. 2. Each email opens with a hook that pulls forward from the previous one. 3. Closing of each email teases the next (curiosity gap, open loop). 4. Final email contains the primary CTA. Earlier emails build trust without asking. 5. Maintain consistent sender voice across the sequence — drip sequences live on continuity.
Cross-channel transformation rules (other directions)
Social → long-form (post-as-seed)
When a social post performs well, expand it. Rules:
1. The post becomes the lede of the long-form piece. 2. Add 3–5 sections that defend or extend the original claim. 3. Add evidence the original post could not carry (data, named cases, citations). 4. Replace the social CTA with a long-form closing (callback or restated stake).
Email → social (newsletter-as-thread)
1. Pick the strongest single argument from the newsletter. 2. Recast as a thread or single post. 3. Preserve the personal frame — newsletter readers expect intimacy; social readers reward authenticity.
Marketing copy → social
Generally avoid. Marketing copy that reads like marketing copy on social produces low engagement. If forced, strip all sales language and focus on a single insight from the marketing piece.
Hooks per channel
Cross-reference samber/cc-skills@copywriting-hooks for the full hook catalog. Per channel grouping, the permitted hook types are:
| Grouping | Permitted hooks |
|---|---|
| Long-form | Scene · contrarian · curiosity gap · concrete detail · stat · question · historical analogy · time anchor · authority |
| Social | Bold claim · direct problem · concrete detail · contrarian · stat · conditional · pattern interrupt |
| Email & newsletter | Personal confession · curiosity gap · open loop · conditional · personal frame |
| Marketing copy | Promise · direct problem · authority · stat · benefit |
CTAs per channel
Cross-reference samber/cc-skills@copywriting-cta for the full CTA archetype catalog. Per channel grouping:
| Grouping | CTA archetypes |
|---|---|
| Long-form | Practical next step · callback · restated stake · transitional asset (lead magnet) |
| Social | Specific reply prompt · link to long-form · profile visit nudge |
| Email & newsletter | Single P.S. CTA · reply prompt · paid-tier tease (if applicable) |
| Marketing copy | Direct action button · book a call · free trial · pricing page |
Anti-patterns per channel
| Channel | Anti-patterns |
|---|---|
| Long-form | Slow scene-setting in technical pieces; closing with generic "What do you think?"; lists-as-article (when prose would do) |
| Social | Hashtag spam; emoji as personality substitute; threading what should be one post; open-ended questions as CTA |
| Email & newsletter | Salesy subject lines; multiple competing CTAs; ignoring preview text; long paragraphs (email clients render them as walls) |
| Marketing copy | "Click here", "Learn more"; feature lists without translation; testimonials without faces and names; multiple CTAs per page |
When ADAPT mode emits a separate file
Two output options for ADAPT mode:
1. Inline channel override section appended to PROSE.md (preferred when channels share most rules) 2. Standalone `PROSE-<grouping>.md` (preferred when a channel has substantial divergence, e.g., a B2B SaaS with a wildly different consumer brand for one product line)
Ask the user which they prefer in ADAPT Phase 2.
Discovery Questions
Full intake battery for BUILD Phase 1. Use AskUserQuestion in 2–3 batches; skip any field already supplied by SOUL.md, TONE.md, or prior conversation. Treat as a kickoff session checklist — a 90-minute interview fills the critical fields; subsequent sessions backfill the rest.
Organized by domain. Mandatory fields are marked [M]; the rest are nice-to-have and improve the guide but do not block it.
1. Brand / entity
- [M] Mission in one sentence (founder's words _and_ marketing's words, side by side)
- [M] Message architecture: 9–12 prioritized attributes (Bloomstein BrandSort or equivalent)
- [M] Brand voice owner operationally (CMO, founder, head of content)
- [M] Brand-to-category posture: conformist · adjacent · challenger · outsider
- Brand age, and whether prose still matches current stage
2. Audience
- [M] Reading age and literacy level (GOV.UK aims at age 9; B2B SaaS commonly at 14–16; expert pubs at 18+)
- [M] Domain expertise: Layperson · Practitioner · Expert
- [M] Cultural context: US · UK · France · EU · global · other
- [M] Language(s) the audience reads in, and fluency
- Professional reading habits (newsletters, publications they trust)
- Patience level: consumer scrolling vs professional researching
3. Existing content
- Top 10 highest-performing pieces of the last 12 months, by channel
- Bottom 10 (failure modes are diagnostic)
- Where voice is consistent, where it drifts
- Which pieces were ghostwritten, agency-produced, or AI-assisted
- What the support / sales team says about how customers describe the brand voice
4. Competitors and category conventions
- 3–5 competitors whose prose is studied (or copied) internally
- Default category register (e.g., enterprise B2B "thought leadership")
- Conventions to conform to vs break, and the cost of breaking each
5. Distribution
- [M] Channels in scope (multiSelect: long-form articles · social posts · email & newsletter · marketing copy)
- Cadence per channel (pieces/month)
- Channel-specific constraints (character limits, SEO requirements, deliverability)
- Typical length per channel
6. Writers and operations
- Who writes, in what mix (employees, freelancers, agencies, ghostwritten principals)
- How briefing is done (templates, briefs, voice notes)
- Review workflow (number of rounds, who has veto)
- Editorial calendar planning horizon
- Tools (Google Docs, Notion, CMS, AI assistants)
7. Content goals
- [M] Per channel, primary KPI: awareness · engagement · lead · signup · retention · advocacy
- How prose changes if KPI is awareness vs conversion vs retention
- Relationship between prose distinctiveness and conversion (sometimes inverse for compliance-heavy categories)
8. Constraints
- Legal: regulated claims, disclaimer requirements, IP/trademark conventions
- Regulatory: industry-specific (finance, health, defense, pharma)
- Compliance: GDPR consent language, accessibility (WCAG 2.2 AA)
- Brand safety: topics that are off-limits
- Confidentiality: what cannot be discussed publicly
9. Cultural context
- [M] Locale of brand HQ vs locale of audience
- [M] Language(s) of operation
- Cultural taboos and sensitivities (regional, religious, political)
- Geopolitical positioning (does the brand take positions on global events?)
10. Evolution
- Expected trajectory in 24 months (geographic expansion, product expansion, repositioning)
- Who decides when the guide is updated, and how often
- What would force a major revision (a rebrand, a pivot, a merger)
11. Author archetype (if SOUL.md missing)
- Primary archetype: journalist · engineer · founder · NGO advocate · politician · consultant · executive · community lead · artist · researcher
- Secondary archetype if hybrid
- The principal's idiolect markers if personal-branding context (filler words, sentence length, preferred connectors, recurring metaphors, banned topics, recurring stories with allowed retelling cadence)
12. Tone of voice (if TONE.md missing)
NN/g four dimensions, position on each spectrum:
- Funny ↔ Serious
- Formal ↔ Casual
- Respectful ↔ Irreverent
- Enthusiastic ↔ Matter-of-fact
A short capture here unblocks Phase 3 codification; recommend producing a full TONE.md via samber/cc-skills@copywriting-tone-of-voice-creator afterward for a brand operating at scale.
The Five Layers of Prose
Two organizing principles before codifying any layer.
Style is content, not decoration. Presentation shapes meaning — a faulty rhythm in a sentence can wreck it as surely as a wrong word. Every syntax choice, every breath point, every clause depth decision is a semantic act, not an aesthetic afterthought.
Concision is the baseline. Every word must justify its inclusion. Not brevity for its own sake — lean sentences carry more authority than padded ones because they never ask the reader to work without payoff.
Codify each layer independently. The layers are orthogonal: a brand can have distinctive lexicon and generic syntax (Liquid Death), or distinctive syntax and generic lexicon (Basecamp). Decide per layer where to conform and where to differentiate.
1. Lexicon
Word-level rules. Deterministic, testable.
1.1 Use / avoid A–Z
A 50–200 entry table. GOV.UK is the gold standard — short entries, alternatives provided, exceptions enumerated.
| Term to use | Why | Term to avoid | Why avoid | Exceptions |
|---|---|---|---|---|
| customer | active subject, dignified | user | reductive outside product docs | "user" OK in product docs |
| help | plain | facilitate | empty Latinate | — |
| build / create | concrete | leverage | empty verb | financial sense permitted |
| use | direct | utilize | wordy | — |
| because | causal | due to the fact that | wordy | — |
| make | concrete | deliver (abstract) | GOV.UK: "only pizzas, post and services are delivered" | concrete senses OK |
Why bother: banning a word without an alternative creates a vacuum writers fill with the next-worst word. Always pair ban with replacement.
1.2 Terminology table
Product names, feature names, capitalization, hyphenation, plural forms. Update on every product launch — a guide that drifts from product reality is ignored by engineers.
1.3 Jargon ladder per channel
| Channel grouping | Specialist terms permitted |
|---|---|
| Long-form articles | Up to 5 per piece, with one-line glosses |
| Social posts | Up to 2 per post, no glosses (link to glossary instead) |
| Email & newsletter | Up to 1 per issue |
| Marketing copy | None unless category requires (compliance, deep-tech) |
1.4 Acronyms
Spell out on first use per page (GOV.UK convention) OR allow known acronyms unexpanded (Microsoft convention). Pick one. The choice depends on audience expertise — Expert audiences resent expansion of acronyms they own.
1.5 Naming conventions
Company, products, features, methodologies. Include possessive forms. Critical for brands with non-standard capitalization (innocent, samber/lo).
1.6 Foreign words
Italics or not, accents preserved or stripped, translation in parentheses or not. Especially critical for French-origin brands writing in English. See multilingual.md.
1.7 Technical depth scale
Three-level: Layperson · Practitioner · Expert. Allocate per channel. Mixing levels within a single piece is the dominant readability failure.
2. Syntax
Sentence- and paragraph-level rules.
2.1 Sentence length distribution
Plain Language Commission default: 15–20 words mean per sentence. Category overrides in category-playbooks.md.
| Metric | Default target | Why |
|---|---|---|
| Mean sentence length | category-specific (10–22) | matches reading age, audience patience |
| Standard deviation | ≥ 6 words per 100-word window | uniform = robotic; variance = human |
| Sentences ≥ 25 words | ≤ 10% | long sentences load short-term memory; rationing protects comprehension |
| Sentences ≤ 8 words | ≥ 15% | punch sentences carry rhythm; their absence = monotony |
Diagnose: 1- run a Python nltk.sent_tokenize + word-count script on a 1000-word sample; 2- Hemingway readability for grade-level cross-check; 3- read the piece aloud, mark every breath point.
2.2 Sentence types
- Declarative as default
- Rhetorical questions: max 1 per long-form piece (or banned entirely — they signal low confidence)
- Imperatives: reserved for CTAs and how-to steps
- Exclamations: rationed per punctuation policy
2.3 Clause depth
Max 2 levels of subordination per sentence. Beyond that, comprehension drops. Coordination preferred for B2C, subordination acceptable for B2B and industry.
2.4 Active vs passive
Active default. Passive permitted for: (a) unknown agent, (b) emphasis on object, (c) impersonal scientific register. Codify the exception list — without it, "active voice always" produces awkward science writing.
Zombie test: append "by zombies" after the verb. If it works, it's passive ("The data was analyzed [by zombies]").
2.5 Parallelism
Enforced in lists, headings, tricolons. Parallel structures must share grammatical category (all noun phrases, or all verb phrases starting with the same tense).
2.6 Paragraph length
- Target 2–5 sentences, max 80 words on web
- One-sentence paragraphs permitted for emphasis but ≤ 15% of paragraphs (else they lose impact)
2.7 Paragraph architecture
Codify one or two house structures. Defaults:
- TES: Topic sentence + 2–3 evidence sentences + implication
- PEEL: Point, Evidence, Explanation, Link
- PAS: Problem, Agitation, Solution (newsletters, email)
- Inverted pyramid: most important first (newsletters, news)
3. Rhythm
How sentences move together. Hardest layer to codify; easiest to audit by reading aloud.
3.1 Cadence
Target sentence-length variance such that σ ≥ 6 words per 100-word window. Lower = robotic / AI tell. See anti-patterns.md.
3.2 Breath points
One short sentence (≤ 8 words) every 3–5 sentences. Why: reading is breathing; missing breath points exhaust the reader's working memory.
3.3 Repetition
- Anaphora (repeated openers): permitted in closings and CTAs only, or banned
- Epistrophe (repeated endings): reserved for signature moments
3.4 Callbacks
If the hook names a person, the closing returns to that person. Why: structural symmetry is a low-cost ethos signal that the piece is built, not generated.
3.5 List patterns
- Prefer prose-then-list (introduce in a sentence) over list-then-prose
- Cap lists at 7 items (Miller's law)
- Parallel grammar enforced
3.6 White space
- Long-form: subheading every 200–350 words
- Social posts: line break every 1–3 sentences
- Email & newsletter: one idea per paragraph
4. Structure
Macrostructures across the piece.
4.1 Openings (hooks)
Cross-ref samber/cc-skills@copywriting-hooks for the full catalog. State 3–5 permitted hook types per channel grouping. Forbid:
- Dictionary-definition openings ("Productivity, defined as...")
- "In today's fast-paced world", "À l'heure du tout-numérique"
- Self-referential openings ("This article will discuss...")
4.2 Closings
Cross-ref samber/cc-skills@copywriting-cta for end-of-article CTA codification. State 2–3 permitted closing types. Forbid: generic "Thanks for reading", "What do you think?" unless signature.
4.3 Transitions
Prefer logical connectors (because, therefore, however, but) over additive (also, moreover, furthermore). Ban "Last but not least."
4.4 Headings
Sentence case, no terminal punctuation, frontloaded with topic noun or active verb, scannable in isolation. GOV.UK rule: "Frontload your headings so that the words that match your users' tasks are at the beginning."
4.5 Subheadings
Parallel structure within an article (all questions, or all noun phrases, not mixed).
4.6 Lists
Bulleted vs numbered policy (numbered only when order matters). Leading sentence required. No nested lists beyond one level.
4.7 Asides
Parentheticals capped at one per 200 words.
4.8 Quotations
Attribution style (name, title, organization; no honorifics on second reference). Minimum quote quality bar: does the quote say something the body cannot?
4.9 Citations and links
Inline, frontloaded link text. Never "click here", "read more", "learn more".
4.10 Blockquotes
Reserved for quotes of 25+ words or for high-emphasis claims.
4.11 Reader positioning (psychic distance)
Gardner's psychic distance continuum describes how close or far the reader sits from the action, the brand, or the subject matter. In brand prose (not fiction), the spectrum runs:
| Distance | Register | Example |
|---|---|---|
| Far | Third-person, external, historical | "Acme was founded in 2005 with a mission to reduce infrastructure costs." |
| Medium-far | Category framing, industry truth | "Most infrastructure teams spend 40% of their time firefighting, not building." |
| Medium-close | Reader-adjacent, hypothetical | "Your team is probably dealing with this right now." |
| Close | Second-person present, internal experience | "You open the dashboard. The number is wrong. Again." |
Why it matters: distance controls emotional temperature. Far establishes authority and context. Close creates empathy and drives conversion. Uncontrolled oscillation reads as schizophrenic; deliberate oscillation creates emotional shape.
Default positions per channel
| Channel grouping | Default | Rationale |
|---|---|---|
| Long-form articles | Medium-far opening → close at key anecdote → far for analysis → close at CTA | Authority frame, then humanity, then rigor, then conversion |
| Social posts | Close hook → medium for argument → close for CTA | Hook must land immediately; brevity requires intimacy |
| Email & newsletter | Medium-close default | Personalization convention; inbox is a private channel |
| Marketing copy | Close default for emotional sections, far for proof/credibility | Conversion copy needs immersion; proof copy needs objectivity |
Shift signals
Moving closer: second-person pronouns ("you", "your"), present tense, sensory detail, internal monologue framing ("You're wondering if…"), short sentences, concrete nouns.
Moving farther: third-person, past tense, statistics, brand history, passive constructions, abstract nouns, long sentences.
Diagnose: scan a 1500-word piece and annotate each paragraph with its distance level (F / MF / MC / C). A flat distribution (all one level) means the piece has no emotional shape. High variance without a pattern means oscillation is accidental. The target is an intentional arc.
5. Voice markers
The small set of repeatable signature elements that make the brand recognizable in a blind test. Catalogue 5–12 markers; more is unenforceable.
| Marker | Definition | Example |
|---|---|---|
| Signature moves | Recurring rhetorical move | Basecamp's contrarian one-sentence paragraph |
| Signoffs | Fixed or templated closing line | Innocent's "Win." |
| Recurring metaphors | 2–3 metaphors the brand owns and reuses | — |
| Idioms | Curated short list (drift is a strong signal of distributed-team decay) | — |
| Taboos | Explicit list of phrases the brand never uses | "leverage", "delve", "in today's fast-paced world" |
| Intentional tics | Deliberately quirky moves (named, rationed) | Innocent's lowercase brand name; Oatly's parenthetical meta-commentary |
Rationing matters. Innocent's puns work because they are rationed. Unrationed quirks collapse into self-parody (the Wackywriting failure mode, per Nick Asbury).
6. Punctuation policy
Declare a position on each. The list is non-negotiable — silence creates drift.
| Mark | Position |
|---|---|
| Em dash | Permitted / banned (declare) — current AI-tell candidate; rationed even when permitted |
| En dash | Ranges only (2024–2026) |
| Semicolon | Permitted in long-form, banned in social and email subject lines |
| Colon | Lists, examples, rephrased restatements; max 1 per paragraph |
| Ellipsis | Banned outside direct quotation (AI tell) |
| Parentheses | Rationed: 1 per 200 words |
| Italics | Foreign words, titles of works, technical first-use emphasis |
| Bold | Scannable phrases in long-form only |
| Quotation marks | Single (UK) or double (US) per locale; consistent |
| Exclamation marks | 1 per 1000 words in long-form; 1 per LinkedIn post; 1 per newsletter |
| Brackets | Editorial insertions in quotations only |
| Hyphens | Compound modifiers before nouns (well-known author); maintain a hyphenated-compound list |
| Oxford comma | Declare yes/no, enforce |
| Capitalization | Sentence case for headings (Microsoft, Mailchimp, IBM Carbon default); title case for proper nouns and product names |
7. Formatting policy
| Element | Rule |
|---|---|
| H1 | One per page |
| H2 | Sections |
| H3 | Sub-sections |
| H4+ | Technical docs only |
| Heading length | ≤ 70 characters |
| Bullets | 3–7 items, parallel grammar, leading sentence required, max 1 level of nesting |
| Numbered lists | Only when sequence matters |
| Code blocks | Language tag mandatory, max 30 lines, prose explanation precedes code |
| Images | Sentence-case captions, descriptive alt text mandatory (WCAG 2.2) |
| Callouts (note, warning, tip) | Rationed: 1 per 800 words in long-form |
| Tables | Only when relationship is two-dimensional |
| Links | Frontloaded link text — never "click here", "learn more", "read more" |
Multilingual Prose
For brands operating across languages (especially EN/FR — the dominant case for France-based operators). The core rule: one PROSE.md per language, not a translated single guide. Codify each language natively; maintain a mapping document of shared pillars and divergent rules.
Why per-language guides
A translated guide propagates the source language's rhythm into the target. French sentence structure produces long sentences with subordinate clauses; English brand prose typically favors shorter sentences. A French→English translation that preserves sentence boundaries reads as labored English. An English→French translation that preserves the original mean sentence length reads as choppy French.
The standard is to retarget, not translate. Translators must follow the target-language guide as if writing fresh.
EN variant declaration
Declare one of: US English · UK English · International English. The choice affects:
| Dimension | US | UK | International |
|---|---|---|---|
| Spelling | -or, -ize, color | -our, -ise, colour | -or, -ize (Microsoft default) |
| Date format | MM/DD/YYYY or "March 5, 2026" | DD/MM/YYYY or "5 March 2026" | YYYY-MM-DD (ISO) |
| Single vs double quotes | Double | Single | Double (more globally readable) |
| Comma in dates | "March 5, 2026" | "5 March 2026" | "2026-03-05" |
| Decimal separator | Period (3.14) | Period (3.14) | Period (3.14) |
| Thousands separator | Comma (1,000) | Comma or space (1,000 / 1 000) | Space (1 000) |
| Time format | 12h with AM/PM | 24h or 12h | 24h |
EN ↔ FR word-policing
French words permitted in English brand text
A small whitelist. Adding French outside this list usually reads as affectation:
- Permitted (no italics, no translation): raison d'être, savoir-faire, joie de vivre, cliché, café, salon, genre, milieu, élite, fiancé(e), résumé (US) / CV (UK), déjà vu, faux pas, par excellence
- Permitted with italics on first use: avant-garde, à la carte, en route, in vino veritas
- Banned without translation: parcours (use "journey" or "path"), dispositif (use "system" or "framework"), enjeu (use "stake" or "issue"), accompagnement (use "support"), démarche (use "approach"), mise en œuvre (use "implementation")
English loan-words accepted in French brand text
French anglicisms cluster around marketing / tech vocabulary. Pick a position:
- Permissive (tech, marketing, consulting brands): le marketing, le briefing, le manager, le pitch, le brainstorming, le storytelling, le timing, le mailing
- Restrictive (cultural institutions, traditional consumer brands, government): replace with French equivalents (le brief → la note, le pitch → la présentation, le marketing → la mercatique [rarely used in practice — fallback to context-specific])
- Banned outright in any French brand text: addressing, deliverables, leverager, prioritiser (use prioriser), implémenter (use mettre en œuvre or réaliser), supporter (use prendre en charge)
False cognates EN ↔ FR
The frequent traps for ghostwriting and translation. Reproducing the false cognate is a single-sentence credibility kill for bilingual readers.
| French word | What writers think it means | What it actually means |
|---|---|---|
| éventuellement | eventually | possibly, perhaps |
| actuellement | actually | currently, right now |
| important | important | often: large, significant in size |
| sensible | sensible | sensitive |
| déception | deception | disappointment |
| location | location | rental |
| librairie | library | bookstore |
| journée | journey | day |
| assister | to assist | to attend |
| supporter | to support | to bear / tolerate / put up with |
| prétendre | to pretend | to claim |
| achever | to achieve | to complete / finish |
| réaliser | to realize | to make / produce / accomplish |
| consister | to consist | "consister à" + verb = to involve doing |
| demander | to demand | to ask |
| habit | habit | clothing (an outfit) |
| chair | chair | flesh |
| coin | coin | corner |
| pain | pain | bread |
| chance | chance | luck |
| sensible (the other way too) | EN word | FR translation: rationnel, raisonnable |
Syntactic transfer rules
The transfer budget corrects rhythm differences between languages.
| Direction | Adjustment |
|---|---|
| FR → EN | Cut 20% of words. French sentences carry more subordinate clauses; English brand prose breaks them. |
| EN → FR | Pad 20% of words. French rewards more developed sentences with explicit connectors. |
| FR → EN | Replace nominal-heavy French ("la mise en œuvre du dispositif") with verbal English ("we deploy the system"). |
| EN → FR | Restore some nominalization for register, especially in B2B and consulting. |
Regionalisms and global English
Declare neutral international English when audiences are global; reserve UK or US idioms for matching audiences. Common traps:
- US-only: "out of the gate", "ballpark figure", "touch base", "rain check"
- UK-only: "knackered", "having a chinwag", "spot on"
- Avoid in international: idiom-heavy sports metaphors (cricket, baseball, American football)
French regional variants
French differs across:
- Hexagonal French (France) — default for France-based brands
- Belgian French — septante / nonante for 70 / 90 (vs soixante-dix / quatre-vingt-dix); some lexical differences
- Swiss French — septante, huitante, nonante; different terminology in retail / banking
- Québécois French — significantly different vocabulary (courriel for email, magasiner for shop); different anglicism policies; English borrowings often resisted strongly
- African French — multiple variants; significant local lexicon
Declare which variant. A French brand expanding into Québec should not assume Hexagonal French passes.
Cultural references
- Safe references (cross-cultural): sports for athletes generally, food and seasons, holidays that are local-relevant only when audience matches
- Forbidden for global audiences: region-specific jokes (US Super Bowl, French baccalauréat), political references, religious holidays as default context
Accessibility and inclusion
- People-first language: "people experiencing X" over "the X"
- Singular they (EN) — established convention since at least 1375, codified by Microsoft, Mailchimp, GOV.UK, Atlassian
- French inclusive writing — declare a position. Three common levels:
1. Conservative: masculine generic ("les développeurs"), historic French Academy default 2. Inclusive parentheses: "les développeurs(euses)" — readable, contested 3. Median point: "les développeur·euse·s" — politically loaded in France; banned in government communication since 2017; accepted in some progressive brand contexts; not all screen readers handle it gracefully
- Bias-free language section: maintain a banned-word list for stigmatizing terms (per Microsoft, Mailchimp templates)
Translation workflow recommendation
When a brand operates in multiple languages:
1. Author in the dominant language (usually the brand's HQ language). 2. Translator-as-writer: never use literal translation as published content. The translator must follow the target-language PROSE.md. 3. Terminology table is bilingual: one canonical term per language, mapped. 4. Idiom and metaphor list per language: do not translate idioms; substitute equivalent register. 5. False-cognate list per language pair: maintained in PROSE.md annex. 6. Channel conventions differ across cultures: LinkedIn convention in France differs from US conventions in formality, paragraph length, first-person use. Codify per locale.
Mapping document
When multiple language guides exist, maintain PROSE-MAPPING.md documenting:
- Shared pillars (the brand's voice principles that hold across languages)
- Divergent rules (where each language guide departs)
- Terminology mappings
- Cross-references for idioms and metaphors
PROSE.md Template
Hybrid format: narrative sections per layer + do/don't tables as an annex. The narrative teaches the _why_ (so writers can handle edge cases); the annex provides the scannable enforcement layer.
Length target: 20–60 pages. Beyond that, writers stop reading. Below that, edge cases proliferate. Siemens reduced their brand guidelines from 2,750 to 250 pages by ruthless deletion — that is the discipline.
Skeleton
# PROSE.md — <Brand Name>
> Version <semver> · Last updated <YYYY-MM-DD> · Owner: <name / role> · Status: <draft | active | deprecated>
>
> Read alongside `TONE.md` (emotional posture) and `SOUL.md` (storyteller archetype). Visual identity lives in `DESIGN.md` and is out of scope here.
## Purpose
200 words: who this guide is for, how to use it, what it does not cover, the relationship to TONE.md and SOUL.md.
## The Prose Pillars
5–8 pillars in the form "We write X, not Y", each with a one-sentence rationale and one example. Pillars must be falsifiable. "We write short sentences with concrete subjects; we avoid abstract nominalizations" passes the test. "We write warmly" does not.
## Voice vs. Tone note
One paragraph adapting Mailchimp's formulation: "You have the same voice all the time, but your tone changes." Voice = consistent (this guide). Tone = situational (TONE.md).
## 1. Lexicon
### 1.1 Use / avoid A–Z
[Narrative paragraph explaining the lexicon's center of gravity — e.g., "Anglo-Saxon verbs over Latinate; concrete nouns over abstract; named over generic."]
[Table: 50–200 entries.]
### 1.2 Terminology
[Product names, feature names, capitalization, plural forms.]
### 1.3 Jargon ladder per channel
[Table: which specialist terms permitted in which channel grouping.]
### 1.4 Acronyms · 1.5 Naming · 1.6 Foreign words · 1.7 Technical depth scale
[As applicable; see five-layers.md for full structure.]
## 2. Syntax
[Narrative: this brand's syntax center of gravity in one paragraph.]
### 2.1 Sentence length distribution
[Mean target ± 2 words. Distribution targets. Category default reasoning.]
### 2.2 Sentence types · 2.3 Clauses · 2.4 Active/passive · 2.5 Parallelism · 2.6 Paragraph length · 2.7 Paragraph architecture
[Each subsection: rule + 1-sentence why + example.]
## 3. Rhythm
[Narrative: what cadence sounds like read aloud.]
### 3.1–3.6 Cadence · Breath points · Repetition · Callbacks · List patterns · White space
## 4. Structure
### 4.1 Openings
[3–5 permitted hook types with example openings from prior brand corpus.] [Forbidden openings list.]
### 4.2 Closings · 4.3 Transitions · 4.4 Headings · 4.5 Subheadings · 4.6 Lists · 4.7 Asides · 4.8 Quotations · 4.9 Citations · 4.10 Blockquotes
## 5. Voice Markers
[5–12 markers with rules of use and rationing.]
### 5.1 Signature moves · 5.2 Signoffs · 5.3 Recurring metaphors · 5.4 Idioms · 5.5 Taboos · 5.6 Intentional tics
## 6. Punctuation Policy
[The full table from five-layers.md, adapted to this brand's positions.]
## 7. Formatting Policy
[Heading hierarchy, lists, code blocks, images, callouts, tables, links.]
## 8. Channel Overrides
[One section per in-scope grouping. Each section: deltas on sentence length, paragraph length, hook types, closing types, formatting, CTA.]
### 8.1 Long-form articles
### 8.2 Social posts
### 8.3 Email & newsletter
### 8.4 Marketing copy
## 9. Cultural & Linguistic Adaptation
[English variant (US/UK/intl); French↔English handling; false cognates; transfer budgets; accessibility/inclusion.]
## 10. Anti-LLM Countermeasures
[Banned lexical tells, structural tells, punctuation defaults. The rules LLMs do not follow by default — that is the durable defense.]
## 11. Sample Bank
### 11.1 Before/after pairs (≥ 10)
For each:
- Rule violated
- Original
- Rewrite
- Rule applied
- Why the rewrite is better (one sentence)
### 11.2 Exemplar pieces (≥ 3, annotated paragraph by paragraph)
### 11.3 Anti-exemplars (≥ 2, de-identified)
### 11.4 Hook bank (30+ approved openings)
### 11.5 Closing bank (15+ approved closings)
### 11.6 Transition bank (20+ approved transition phrases)
## 12. Ghostwriting Addendum (per principal, if applicable)
[Per-principal idiolect: 10 signature openings, 5 signature closings, 3 recurring stories with allowed retelling cadence, list of banned topics, list of preferred connectors, average post length, line-break convention.]
---
## Annex A — Do/Don't Quick Reference
The scannable layer. One table per layer; writers can audit a draft in 10 minutes.
### A.1 Lexicon
| ✅ Do | ❌ Don't |
| --- | --- |
| Use plain Anglo-Saxon verbs | Use empty Latinate verbs (leverage, facilitate, utilize) |
| Spell out acronyms on first use | Assume all readers know the acronym |
| Use product names exactly as registered | Improvise capitalization |
### A.2 Syntax
| ✅ Do | ❌ Don't |
| --- | --- |
| Vary sentence length (σ ≥ 6 words per 100-word window) | Write uniformly long or uniformly short sentences |
| Active voice as default | Use passive without a documented reason |
| Cap subordination at 2 levels | Stack subordinate clauses |
### A.3 Rhythm
| ✅ Do | ❌ Don't |
| --- | --- |
| Place a breath sentence (≤ 8 words) every 3–5 sentences | Forget to breathe |
| Use parallel structure in lists and tricolons | Mix grammatical categories in a list |
| Cap lists at 7 items | Write 12-item lists with no grouping |
### A.4 Structure
| ✅ Do | ❌ Don't |
| --- | --- |
| Front-load headings with the topic noun | Open with throat-clearing ("Introduction to...") |
| Use logical connectors (because, therefore, however) | Use additive filler (also, moreover, furthermore, last but not least) |
| Frontloaded link text | "click here", "learn more", "read more" |
### A.5 Voice markers
| ✅ Do | ❌ Don't |
| --- | --- |
| Deploy signature moves at the declared rate | Overuse a signature move into self-parody |
| Maintain the taboo list | Drift into category-default phrasings |
### A.6 Punctuation
| ✅ Do | ❌ Don't |
| --- | --- |
| Enforce the Oxford comma decision consistently | Switch within a piece |
| Ration exclamation marks per the policy | Use exclamations as enthusiasm performance |
| Banned em dash → use comma, colon, parens, or period (if banned) | Keep em dashes when policy says no |
### A.7 Channel overrides
| Channel | Mean sentence length | Paragraph length | Hook style | CTA |
| --- | --- | --- | --- | --- |
| Long-form | 14–18 | 2–5 sentences | Scene · contrarian · stat · concrete detail | Practical next step / callback |
| Social | 8–12 | 1–3 sentences | Bold claim · direct problem · concrete detail | Specific reply prompt |
| Email | 10–14 | 1–2 sentences | Personal frame · curiosity gap | Single primary CTA in P.S. |
| Marketing copy | 8–12 | 1–3 sentences | Promise · direct problem · authority | Direct action button |
---
## Changelog
| Date | Version | Change | Author |
| ---------- | ------- | ------------- | ------ |
| YYYY-MM-DD | 1.0.0 | Initial guide | name |Notes on populating the template
- Pillars are mandatory. Without falsifiable pillars, writers default to invented rules.
- Sample bank is the most-read section. Lead with it in onboarding; treat it as the front door, not the appendix.
- Annex tables are co-located by layer. Editors read top-down narrative for understanding; writers spot-check via the annex on every piece.
- The changelog is part of the trust. A guide updated visibly is a guide writers trust.
Related skills
How it compares
Choose copywriting-prose-creator over general copywriting skills when you need a durable PROSE.md mechanics spec separate from tone and channel-specific adaptation.
FAQ
What is copywriting-prose-creator?
Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature moves) independent of emotional tone. Output: PROSE.md. Three modes: BUILD a
When should I use copywriting-prose-creator?
Codifies how someone or a brand writes — prose mechanics (lexicon, syntax, rhythm, structure, signature moves) independent of emotional tone. Output: PROSE.md. Three modes: BUILD a
Is copywriting-prose-creator safe to install?
Review the Security Audits panel on this page before production use.