
Copy Editing
- 554 installs
- 23.5k repo stars
- Updated July 17, 2026
- alirezarezvani/claude-skills
copy-editing is a Claude Code skill that detects AI-flat prose using burstiness and vocabulary metrics via ai_content_detector.py, then humanizes landing pages, emails, and docs before publish.
About
Copy Editing is an agent skill that helps developers polish marketing and product copy so it reads human rather than model-flat. It explains how burstiness—the variance of sentence lengths—flags overly uniform AI rhythm, with clear CV bands from natural through strong AI signals. It also covers vocabulary diversity via type-token ratio across sliding windows, where repetitive safe word choices inflate AI likelihood. For each flagged pattern, the skill gives concrete rewrite moves: insert short sentences, split long lines, allow fragments, and alternate paragraph sizes. It is framed as reference material for an AI content detector workflow, so your agent can review drafts before lifecycle emails, blog posts, or landing hero sections go live. Use it when generated copy sounds generic, when you worry about detection or reader trust, or when you want consistent editorial pass rules the agent can apply pass after pass.
- Burstiness scoring via sentence-length coefficient of variation with interpreted AI probability bands
- Vocabulary diversity checks using type-token ratio in 200-word sliding windows
- Actionable humanization fixes: short punches, fragments, and deliberate length variance
- Reference aligned to ai_content_detector.py detection methods
Copy Editing by the numbers
- 554 all-time installs (skills.sh)
- +10 installs in the week ending Jun 22, 2026 (Skillselion tracking)
- Ranked #714 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Security screen: LOW risk (skills.sh audit)
- Data as of Jul 31, 2026 (Skillselion catalog sync)
npx skills add https://github.com/alirezarezvani/claude-skills --skill copy-editingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 554 |
|---|---|
| repo stars | ★ 23.5k |
| Security audit | 3 / 3 scanners passed |
| Last updated | July 17, 2026 |
| Repository | alirezarezvani/claude-skills ↗ |
How do you detect and fix AI-flat writing?
Detect AI-flat prose with burstiness and vocabulary metrics, then humanize landing pages, emails, and docs before publish.
Who is it for?
Developers publishing landing pages, emails, or docs who need metric-backed detection and rewriting of AI-flat prose before launch.
Skip if: Legal compliance review, translation workflows, or code refactoring tasks unrelated to marketing or documentation copy quality.
When should I use this skill?
The user needs to humanize AI-generated copy, check burstiness scores, or polish landing pages and emails before publishing.
What you get
Humanized copy drafts, burstiness CV scores, AI probability ratings, and publish-ready landing page or email text.
- Humanized copy
- Burstiness and AI probability report
By the numbers
- ai_content_detector.py implements 3 detection methods including burstiness CV scoring
- CV 0.50+ maps to 0–30% AI probability; CV 0.35–0.49 maps to 30–50% AI probability
Files
Copy Editing
You are an expert copy editor specializing in marketing and conversion copy. Your goal is to systematically improve existing copy through focused editing passes while preserving the core message.
Core Philosophy
Check for product marketing context first: If .claude/product-marketing-context.md exists, read it before editing. Use brand voice and customer language from that context to guide your edits.
Good copy editing isn't about rewriting—it's about enhancing. Each pass focuses on one dimension, catching issues that get missed when you try to fix everything at once.
Key principles:
- Don't change the core message; focus on enhancing it
- Multiple focused passes beat one unfocused review
- Each edit should have a clear reason
- Preserve the author's voice while improving clarity
---
The Seven Sweeps Framework
Edit copy through seven sequential passes, each focusing on one dimension. After each sweep, loop back to check previous sweeps aren't compromised.
Sweep 1: Clarity
Focus: Can the reader understand what you're saying?
What to check:
- Confusing sentence structures
- Unclear pronoun references
- Jargon or insider language
- Ambiguous statements
- Missing context
Common clarity killers:
- Sentences trying to say too much
- Abstract language instead of concrete
- Assuming reader knowledge they don't have
- Burying the point in qualifications
Process: 1. Score the draft mechanically first: python3 scripts/readability_scorer.py --file draft.md (Flesch score, passive-voice %, filler-word count; add --json for pipelines). Anything it flags is your starting highlight list. 2. Read through quickly, highlighting unclear parts the scorer can't see 3. Don't correct yet—just note problem areas 4. After marking issues, recommend specific edits 5. Verify edits maintain the original intent — re-run the scorer; the Flesch score should improve, not regress
After this sweep: Confirm the "Rule of One" (one main idea per section) and "You Rule" (copy speaks to the reader) are intact.
---
Sweep 2: Voice and Tone
Focus: Is the copy consistent in how it sounds?
What to check:
- Shifts between formal and casual
- Inconsistent brand personality
- Mood changes that feel jarring
- Word choices that don't match the brand
Common voice issues:
- Starting casual, becoming corporate
- Mixing "we" and "the company" references
- Humor in some places, serious in others (unintentionally)
- Technical language appearing randomly
Process: 1. Read aloud to hear inconsistencies 2. Mark where tone shifts unexpectedly 3. Recommend edits that smooth transitions 4. Ensure personality remains throughout
After this sweep: Return to Clarity Sweep to ensure voice edits didn't introduce confusion.
---
Sweep 3: So What
Focus: Does every claim answer "why should I care?"
What to check:
- Features without benefits
- Claims without consequences
- Statements that don't connect to reader's life
- Missing "which means..." bridges
The So What test: For every statement, ask "Okay, so what?" If the copy doesn't answer that question with a deeper benefit, it needs work.
❌ "Our platform uses AI-powered analytics" So what? ✅ "Our AI-powered analytics surface insights you'd miss manually—so you can make better decisions in half the time"
Common So What failures:
- Feature lists without benefit connections
- Impressive-sounding claims that don't land
- Technical capabilities without outcomes
- Company achievements that don't help the reader
Process: 1. Read each claim and literally ask "so what?" 2. Highlight claims missing the answer 3. Add the benefit bridge or deeper meaning 4. Ensure benefits connect to real reader desires
After this sweep: Return to Voice and Tone, then Clarity.
---
Sweep 4: Prove It
Focus: Is every claim supported with evidence?
What to check:
- Unsubstantiated claims
- Missing social proof
- Assertions without backup
- "Best" or "leading" without evidence
Types of proof to look for:
- Testimonials with names and specifics
- Case study references
- Statistics and data
- Third-party validation
- Guarantees and risk reversals
- Customer logos
- Review scores
Common proof gaps:
- "Trusted by thousands" (which thousands?)
- "Industry-leading" (according to whom?)
- "Customers love us" (show them saying it)
- Results claims without specifics
Process: 1. Identify every claim that needs proof 2. Check if proof exists nearby 3. Flag unsupported assertions 4. Recommend adding proof or softening claims
After this sweep: Return to So What, Voice and Tone, then Clarity.
---
Sweep 5: Specificity
Focus: Is the copy concrete enough to be compelling?
What to check:
- Vague language ("improve," "enhance," "optimize")
- Generic statements that could apply to anyone
- Round numbers that feel made up
- Missing details that would make it real
Specificity upgrades:
| Vague | Specific |
|---|---|
| Save time | Save 4 hours every week |
| Many customers | 2,847 teams |
| Fast results | Results in 14 days |
| Improve your workflow | Cut your reporting time in half |
| Great support | Response within 2 hours |
Common specificity issues:
- Adjectives doing the work nouns should do
- Benefits without quantification
- Outcomes without timeframes
- Claims without concrete examples
Process: 1. Highlight vague words and phrases 2. Ask "Can this be more specific?" 3. Add numbers, timeframes, or examples 4. Remove content that can't be made specific (it's probably filler)
After this sweep: Return to Prove It, So What, Voice and Tone, then Clarity.
---
Sweep 6: Heightened Emotion
Focus: Does the copy make the reader feel something?
What to check:
- Flat, informational language
- Missing emotional triggers
- Pain points mentioned but not felt
- Aspirations stated but not evoked
Emotional dimensions to consider:
- Pain of the current state
- Frustration with alternatives
- Fear of missing out
- Desire for transformation
- Pride in making smart choices
- Relief from solving the problem
Techniques for heightening emotion:
- Paint the "before" state vividly
- Use sensory language
- Tell micro-stories
- Reference shared experiences
- Ask questions that prompt reflection
Process: 1. Read for emotional impact—does it move you? 2. Identify flat sections that should resonate 3. Add emotional texture while staying authentic 4. Ensure emotion serves the message (not manipulation)
After this sweep: Return to Specificity, Prove It, So What, Voice and Tone, then Clarity.
---
Sweep 7: Zero Risk
Focus: Have we removed every barrier to action?
What to check:
- Friction near CTAs
- Unanswered objections
- Missing trust signals
- Unclear next steps
- Hidden costs or surprises
Risk reducers to look for:
- Money-back guarantees
- Free trials
- "No credit card required"
- "Cancel anytime"
- Social proof near CTA
- Clear expectations of what happens next
- Privacy assurances
Common risk issues:
- CTA asks for commitment without earning trust
- Objections raised but not addressed
- Fine print that creates doubt
- Vague "Contact us" instead of clear next step
Process: 1. Focus on sections near CTAs 2. List every reason someone might hesitate 3. Check if the copy addresses each concern 4. Add risk reversals or trust signals as needed
After this sweep: Return through all previous sweeps one final time: Heightened Emotion, Specificity, Prove It, So What, Voice and Tone, Clarity.
---
Quick-Pass Editing Checks
Use these for faster reviews when a full seven-sweep process isn't needed.
AI-Pattern Check
If the draft may be AI-generated (or AI-assisted), run the detector before editing:
python3 scripts/ai_content_detector.py draft.md --json # no arg = --demo modeIt scores burstiness, vocabulary diversity, and stock-phrase density. A high AI-likelihood score means the piece needs content-humanizer treatment before copy editing — polishing AI mush produces polished AI mush.
Word-Level Checks
Cut these words:
- Very, really, extremely, incredibly (weak intensifiers)
- Just, actually, basically (filler)
- In order to (use "to")
- That (often unnecessary)
- Things, stuff (vague)
Replace these:
| Weak | Strong |
|---|---|
| Utilize | Use |
| Implement | Set up |
| Leverage | Use |
| Facilitate | Help |
| Innovative | New |
| Robust | Strong |
| Seamless | Smooth |
| Cutting-edge | New/Modern |
Watch for:
- Adverbs (usually unnecessary)
- Passive voice (switch to active)
- Nominalizations (verb → noun: "make a decision" → "decide")
Sentence-Level Checks
- One idea per sentence
- Vary sentence length (mix short and long)
- Front-load important information
- Max 3 conjunctions per sentence
- No more than 25 words (usually)
Paragraph-Level Checks
- One topic per paragraph
- Short paragraphs (2-4 sentences for web)
- Strong opening sentences
- Logical flow between paragraphs
- White space for scannability
---
Copy Editing Checklist
Before You Start
- [ ] Understand the goal of this copy
- [ ] Know the target audience
- [ ] Identify the desired action
- [ ] Read through once without editing
Clarity (Sweep 1)
- [ ] Every sentence is immediately understandable
- [ ] No jargon without explanation
- [ ] Pronouns have clear references
- [ ] No sentences trying to do too much
Voice & Tone (Sweep 2)
- [ ] Consistent formality level throughout
- [ ] Brand personality maintained
- [ ] No jarring shifts in mood
- [ ] Reads well aloud
So What (Sweep 3)
- [ ] Every feature connects to a benefit
- [ ] Claims answer "why should I care?"
- [ ] Benefits connect to real desires
- [ ] No impressive-but-empty statements
Prove It (Sweep 4)
- [ ] Claims are substantiated
- [ ] Social proof is specific and attributed
- [ ] Numbers and stats have sources
- [ ] No unearned superlatives
Specificity (Sweep 5)
- [ ] Vague words replaced with concrete ones
- [ ] Numbers and timeframes included
- [ ] Generic statements made specific
- [ ] Filler content removed
Heightened Emotion (Sweep 6)
- [ ] Copy evokes feeling, not just information
- [ ] Pain points feel real
- [ ] Aspirations feel achievable
- [ ] Emotion serves the message authentically
Zero Risk (Sweep 7)
- [ ] Objections addressed near CTA
- [ ] Trust signals present
- [ ] Next steps are crystal clear
- [ ] Risk reversals stated (guarantee, trial, etc.)
Final Checks
- [ ] No typos or grammatical errors
- [ ] Consistent formatting
- [ ] Links work (if applicable)
- [ ] Core message preserved through all edits
---
Common Copy Problems & Fixes
Problem: Wall of Features
Symptom: List of what the product does without why it matters Fix: Add "which means..." after each feature to bridge to benefits
Problem: Corporate Speak
Symptom: "Leverage synergies to optimize outcomes" Fix: Ask "How would a human say this?" and use those words
Problem: Weak Opening
Symptom: Starting with company history or vague statements Fix: Lead with the reader's problem or desired outcome
Problem: Buried CTA
Symptom: The ask comes after too much buildup, or isn't clear Fix: Make the CTA obvious, early, and repeated
Problem: No Proof
Symptom: "Customers love us" with no evidence Fix: Add specific testimonials, numbers, or case references
Problem: Generic Claims
Symptom: "We help businesses grow" Fix: Specify who, how, and by how much
Problem: Mixed Audiences
Symptom: Copy tries to speak to everyone, resonates with no one Fix: Pick one audience and write directly to them
Problem: Feature Overload
Symptom: Listing every capability, overwhelming the reader Fix: Focus on 3-5 key benefits that matter most to the audience
---
Working with Copy Sweeps
When editing collaboratively:
1. Run a sweep and present findings - Show what you found, why it's an issue 2. Recommend specific edits - Don't just identify problems; propose solutions 3. Request the updated copy - Let the author make final decisions 4. Verify previous sweeps - After each round of edits, re-check earlier sweeps 5. Repeat until clean - Continue until a full sweep finds no new issues
This iterative process ensures each edit doesn't create new problems while respecting the author's ownership of the copy.
---
References
- Plain English Alternatives: Replace complex words with simpler alternatives
---
Task-Specific Questions
1. What's the goal of this copy? (Awareness, conversion, retention) 2. What action should readers take? 3. Are there specific concerns or known issues? 4. What proof/evidence do you have available?
---
When to Use Each Skill
| Task | Skill to Use |
|---|---|
| Writing new page copy from scratch | copywriting |
| Reviewing and improving existing copy | copy-editing (this skill) |
| Editing copy you just wrote | copy-editing (this skill) |
| Structural or strategic page changes | page-cro |
---
Proactive Triggers
Surface these issues WITHOUT being asked when you notice them in context:
- Copy is submitted for editing without a stated goal → Ask for the target action and audience before starting any sweeps; editing without this context guarantees misaligned feedback.
- Multiple tone shifts detected → Flag Sweep 2 failure immediately; note the specific lines where voice breaks and propose fixes before continuing.
- Features outnumber benefits 2:1 or more → Raise the "So What" alarm early in the review; this is the single most common conversion killer.
- Superlatives without proof ("best," "leading," "most trusted") → Flag each instance in Sweep 4 and request the evidence or softer language alternatives.
- CTA is vague or buried → Call this out in Sweep 7 before delivering any other feedback — it's the highest-impact fix.
---
Output Artifacts
| When you ask for... | You get... |
|---|---|
| A full copy review | Seven-sweep structured report with specific issues, proposed edits, and rationale for each change |
| A quick copy pass | Word- and sentence-level edits with tracked-change style annotations |
| A copy editing checklist run | Completed checklist with pass/fail per section and priority fixes |
| Specific sweep only (e.g., Clarity) | Focused report for that sweep with before/after examples |
| Final polish | Clean edited version of the copy with a summary of all changes made |
---
Communication
All output follows the structured communication standard:
- Bottom line first — state the overall copy health before diving into issues
- What + Why + How — every flagged issue gets: what's wrong, why it hurts conversion, how to fix it
- Edits have reasons — never change words without explaining the principle
- Confidence tagging — 🟢 clear improvement / 🟡 judgment call / 🔴 needs author input
Deliver findings sweep-by-sweep. Don't dump all issues at once. Prioritize by conversion impact, not writing preference.
---
Related Skills
- marketing-context: USE as foundation before editing — provides brand voice, ICP, and tone benchmarks. NOT a substitute for reading the copy itself.
- copywriting: USE when the copy needs to be rewritten from scratch rather than edited. NOT for polishing existing drafts.
- content-strategy: USE when the problem is what to say, not how to say it. NOT for line-level improvements.
- social-content: USE when edited copy needs to be adapted for social platforms. NOT for page-level editing.
- marketing-ideas: USE when the client needs a new marketing angle entirely. NOT for editorial improvement.
- content-humanizer: USE when AI-generated copy needs to pass the human test before copy editing begins. NOT for structural review.
- ab-test-setup: USE when disagreement on copy variants needs data to resolve. NOT for the editing process itself.
AI Content Detection Patterns
Reference for ai_content_detector.py. Explains the three detection methods and how to humanize flagged content.
Method 1: Burstiness (sentence length variance)
What it measures: The coefficient of variation (CV) of sentence lengths across the text.
Why it works: Human writers naturally vary between short punchy sentences (4-8 words) and longer explanatory ones (20-35 words). AI tends to produce consistently medium-length sentences (12-20 words), creating a "flat" rhythm.
| CV Range | Interpretation | AI Probability |
|---|---|---|
| 0.50+ | High variance — natural human rhythm | 0-30% |
| 0.35-0.49 | Moderate variance — could be either | 30-50% |
| 0.20-0.34 | Low variance — suspiciously uniform | 50-80% |
| < 0.20 | Very flat — strong AI signal | 80-100% |
How to fix flagged text:
- Deliberately insert short sentences. "This matters." "Here's why."
- Break one long sentence into two short ones, then follow with a 25+ word sentence
- Use fragments where tone allows. "Not always. But often enough."
- Vary paragraph length too: alternate 1-sentence and 3-4 sentence paragraphs
Method 2: Vocabulary diversity (Type-Token Ratio)
What it measures: The ratio of unique words to total words in sliding 200-word windows.
Why it works: AI models tend to reuse the same "safe" vocabulary — common verbs, generic adjectives, standard connectors. Human writers use more domain-specific terminology, colloquialisms, and varied word choices.
| TTR Range | Interpretation | AI Probability |
|---|---|---|
| 0.60+ | Rich vocabulary — likely human | 0-25% |
| 0.45-0.59 | Average — could be either | 25-50% |
| 0.35-0.44 | Repetitive — AI-like | 50-75% |
| < 0.35 | Very repetitive — strong AI signal | 75-100% |
How to fix flagged text:
- Replace generic verbs ("use", "make", "get") with specific ones ("wield", "craft", "extract")
- Add domain jargon where your audience expects it
- Vary connectors: instead of always "however", try "still", "yet", "that said", "then again"
- Remove filler phrases that inflate word count without adding meaning
Method 3: Known AI phrases
30 phrases that appear disproportionately in LLM-generated text. These aren't wrong individually — some are perfectly fine in context — but high density signals AI origin.
Density thresholds:
| Density (per 1K words) | Interpretation |
|---|---|
| 0-2 | Normal range |
| 3-5 | Elevated — review flagged phrases |
| 6+ | High — likely AI-generated |
The 10 most common AI phrases to watch:
1. "In today's digital landscape" — replace with specific context 2. "It's worth noting that" — just state the fact 3. "Leverage" — use "use" unless specifically about financial leverage 4. "Delve into" — use "explore", "examine", or "look at" 5. "Game-changer" — use a specific description of impact 6. "Comprehensive guide" — be specific about what's covered 7. "Seamlessly integrate" — describe the actual integration 8. "Robust solution" — describe what makes it robust 9. "Cutting-edge" — name the specific advancement 10. "Empower you to" — just say what it enables
The fix: Replace generic AI phrases with specific, concrete language. "In today's digital landscape" → "Since Google's March 2025 core update". "Leverage AI tools" → "Use GPT-4 for first-draft outlines".
Composite scoring
Composite = Burstiness × 0.35 + Vocabulary × 0.30 + Phrases × 0.35
0-20: LIKELY_HUMAN — no action needed
21-50: MIXED — review flagged passages, humanize selectively
51-100: LIKELY_AI — significant rewriting recommendedImportant caveats
- This is heuristic, not proof. Technical documentation often has low burstiness and TTR naturally.
- Some AI phrases are perfectly appropriate in context. Don't mechanically remove them all.
- The goal is to make content SOUND human, not to prove it IS human.
- Run this tool AFTER writing, not during — it's an editing pass, not a writing constraint.
Plain English Alternatives
Replace complex or pompous words with plain English alternatives.
Source: Plain English Campaign A-Z of Alternative Words (2001), Australian Government Style Manual (2024), plainlanguage.gov
---
A
| Complex | Plain Alternative |
|---|---|
| (an) absence of | no, none |
| abundance | enough, plenty, many |
| accede to | allow, agree to |
| accelerate | speed up |
| accommodate | meet, hold, house |
| accomplish | do, finish, complete |
| accordingly | so, therefore |
| acknowledge | thank you for, confirm |
| acquire | get, buy, obtain |
| additional | extra, more |
| adjacent | next to |
| advantageous | useful, helpful |
| advise | tell, say, inform |
| aforesaid | this, earlier |
| aggregate | total |
| alleviate | ease, reduce |
| allocate | give, share, assign |
| alternative | other, choice |
| ameliorate | improve |
| anticipate | expect |
| apparent | clear, obvious |
| appreciable | large, noticeable |
| appropriate | proper, right, suitable |
| approximately | about, roughly |
| ascertain | find out |
| assistance | help |
| at the present time | now |
| attempt | try |
| authorise | allow, let |
---
B
| Complex | Plain Alternative |
|---|---|
| belated | late |
| beneficial | helpful, useful |
| bestow | give |
| by means of | by |
---
C
| Complex | Plain Alternative |
|---|---|
| calculate | work out |
| cease | stop, end |
| circumvent | avoid, get around |
| clarification | explanation |
| commence | start, begin |
| communicate | tell, talk, write |
| competent | able |
| compile | collect, make |
| complete | fill in, finish |
| component | part |
| comprise | include, make up |
| (it is) compulsory | (you) must |
| conceal | hide |
| concerning | about |
| consequently | so |
| considerable | large, great, much |
| constitute | make up, form |
| consult | ask, talk to |
| consumption | use |
| currently | now |
---
D
| Complex | Plain Alternative |
|---|---|
| deduct | take off |
| deem | treat as, consider |
| defer | delay, put off |
| deficiency | lack |
| delete | remove, cross out |
| demonstrate | show, prove |
| denote | show, mean |
| designate | name, appoint |
| despatch/dispatch | send |
| determine | decide, find out |
| detrimental | harmful |
| diminish | reduce, lessen |
| discontinue | stop |
| disseminate | spread, distribute |
| documentation | papers, documents |
| due to the fact that | because |
| duration | time, length |
| dwelling | home |
---
E
| Complex | Plain Alternative |
|---|---|
| economical | cheap, good value |
| eligible | allowed, qualified |
| elucidate | explain |
| enable | allow |
| encounter | meet |
| endeavour | try |
| enquire | ask |
| ensure | make sure |
| entitlement | right |
| envisage | expect |
| equivalent | equal, the same |
| erroneous | wrong |
| establish | set up, show |
| evaluate | assess, test |
| excessive | too much |
| exclusively | only |
| exempt | free from |
| expedite | speed up |
| expenditure | spending |
| expire | run out |
---
F
| Complex | Plain Alternative |
|---|---|
| fabricate | make |
| facilitate | help, make possible |
| finalise | finish, complete |
| following | after |
| for the purpose of | to, for |
| for the reason that | because |
| forthwith | now, at once |
| forward | send |
| frequently | often |
| furnish | give, provide |
| furthermore | also, and |
---
G-H
| Complex | Plain Alternative |
|---|---|
| generate | produce, create |
| henceforth | from now on |
| hitherto | until now |
---
I
| Complex | Plain Alternative |
|---|---|
| if and when | if, when |
| illustrate | show |
| immediately | at once, now |
| implement | carry out, do |
| imply | suggest |
| in accordance with | under, following |
| in addition to | and, also |
| in conjunction with | with |
| in excess of | more than |
| in lieu of | instead of |
| in order to | to |
| in receipt of | receive |
| in relation to | about |
| in respect of | about, for |
| in the event of | if |
| in the majority of instances | most, usually |
| in the near future | soon |
| in view of the fact that | because |
| inception | start |
| indicate | show, suggest |
| inform | tell |
| initiate | start, begin |
| insert | put in |
| instances | cases |
| irrespective of | despite |
| issue | give, send |
---
L-M
| Complex | Plain Alternative |
|---|---|
| (a) large number of | many |
| liaise with | work with, talk to |
| locality | place, area |
| locate | find |
| magnitude | size |
| (it is) mandatory | (you) must |
| manner | way |
| modification | change |
| moreover | also, and |
---
N-O
| Complex | Plain Alternative |
|---|---|
| negligible | small |
| nevertheless | but, however |
| notify | tell |
| notwithstanding | despite, even if |
| numerous | many |
| objective | aim, goal |
| (it is) obligatory | (you) must |
| obtain | get |
| occasioned by | caused by |
| on behalf of | for |
| on numerous occasions | often |
| on receipt of | when you get |
| on the grounds that | because |
| operate | work, run |
| optimum | best |
| option | choice |
| otherwise | or |
| outstanding | unpaid |
| owing to | because |
---
P
| Complex | Plain Alternative |
|---|---|
| partially | partly |
| participate | take part |
| particulars | details |
| per annum | a year |
| perform | do |
| permit | let, allow |
| personnel | staff, people |
| peruse | read |
| possess | have, own |
| practically | almost |
| predominant | main |
| prescribe | set |
| preserve | keep |
| previous | earlier, before |
| principal | main |
| prior to | before |
| proceed | go ahead |
| procure | get |
| prohibit | ban, stop |
| promptly | quickly |
| provide | give |
| provided that | if |
| provisions | rules, terms |
| proximity | nearness |
| purchase | buy |
| pursuant to | under |
---
R
| Complex | Plain Alternative |
|---|---|
| reconsider | think again |
| reduction | cut |
| referred to as | called |
| regarding | about |
| reimburse | repay |
| reiterate | repeat |
| relating to | about |
| remain | stay |
| remainder | rest |
| remuneration | pay |
| render | make, give |
| represent | stand for |
| request | ask |
| require | need |
| residence | home |
| retain | keep |
| revised | changed, new |
---
S
| Complex | Plain Alternative |
|---|---|
| scrutinise | examine, check |
| select | choose |
| solely | only |
| specified | given, stated |
| state | say |
| statutory | legal, by law |
| subject to | depending on |
| submit | send, give |
| subsequent to | after |
| subsequently | later |
| substantial | large, much |
| sufficient | enough |
| supplement | add to |
| supplementary | extra |
---
T-U
| Complex | Plain Alternative |
|---|---|
| terminate | end, stop |
| thereafter | then |
| thereby | by this |
| thus | so |
| to date | so far |
| transfer | move |
| transmit | send |
| ultimately | in the end |
| undertake | agree, do |
| uniform | same |
| utilise | use |
---
V-Z
| Complex | Plain Alternative |
|---|---|
| variation | change |
| virtually | almost |
| visualise | imagine, see |
| ways and means | ways |
| whatsoever | any |
| with a view to | to |
| with effect from | from |
| with reference to | about |
| with regard to | about |
| with respect to | about |
| zone | area |
---
Phrases to Remove Entirely
These phrases often add nothing. Delete them:
- a total of
- absolutely
- actually
- all things being equal
- as a matter of fact
- at the end of the day
- at this moment in time
- basically
- currently (when "now" or nothing works)
- I am of the opinion that (use: I think)
- in due course (use: soon, or say when)
- in the final analysis
- it should be understood
- last but not least
- obviously
- of course
- quite
- really
- the fact of the matter is
- to all intents and purposes
- very
#!/usr/bin/env python3
"""
ai_content_detector.py — Detect AI-generated content patterns.
Three detection methods:
1. Burstiness analysis — human writing has high variance in sentence length;
AI writing has consistently medium-length sentences
2. Vocabulary diversity (Type-Token Ratio) — AI reuses words more than humans
3. Known AI phrases — specific phrases that appear disproportionately in
AI-generated text
This is a heuristic tool, NOT a proof engine. False positives are expected.
The goal is to flag passages that FEEL AI-generated so a human editor can
inject voice, variance, and specificity.
Usage:
python ai_content_detector.py article.md
python ai_content_detector.py article.md --json
python ai_content_detector.py --demo
Scoring:
0-20 = likely human (high burstiness, diverse vocab, no AI phrases)
21-50 = mixed signals (review flagged passages)
51-100 = likely AI (flat burstiness, repetitive vocab, AI phrase density)
"""
from __future__ import annotations
import argparse
import json
import math
import re
import sys
from collections import Counter
from pathlib import Path
# --- Known AI phrases (commonly overrepresented in LLM output) ---
AI_PHRASES = [
"in today's digital landscape",
"in today's fast-paced",
"it's worth noting that",
"it is important to note",
"delve into",
"dive deep into",
"leverage",
"game-changer",
"game changer",
"unlock the potential",
"unlock the power",
"harness the power",
"at the end of the day",
"in conclusion",
"in summary",
"navigating the complexities",
"a comprehensive guide",
"seamlessly integrate",
"robust solution",
"cutting-edge",
"state-of-the-art",
"empower you to",
"take your .* to the next level",
"in the realm of",
"tapestry of",
"multifaceted",
"it's crucial to",
"paramount",
"foster a .* environment",
"elevate your",
]
AI_PHRASE_RES = [re.compile(p, re.IGNORECASE) for p in AI_PHRASES]
# --- Sentence splitting ---
SENTENCE_RE = re.compile(r"[^.!?]+[.!?]+", re.DOTALL)
WORD_RE = re.compile(r"[a-zA-Z]+")
DEMO_CONTENT = """In today's digital landscape, leveraging AI tools has become a game-changer for content creators. It's worth noting that the ability to harness the power of large language models can unlock the potential of your marketing efforts. This comprehensive guide will delve into the multifaceted world of AI-assisted content creation.
The integration of AI into content workflows is a robust solution that seamlessly integrates with existing processes. By navigating the complexities of modern content production, you can elevate your brand's voice and foster a creative environment that empowers your team to take their content to the next level.
At the end of the day, it's crucial to understand that AI is a tool, not a replacement. The cutting-edge capabilities of state-of-the-art models are paramount for staying competitive in the realm of digital marketing. In conclusion, the tapestry of modern content creation requires both human creativity and artificial intelligence working in harmony."""
def extract_sentences(text):
body = re.sub(r"^---.*?---\s*", "", text, count=1, flags=re.DOTALL)
body = re.sub(r"^#+\s+.*$", "", body, flags=re.MULTILINE)
body = re.sub(r"```.*?```", "", body, flags=re.DOTALL)
body = re.sub(r"`[^`]+`", "", body)
body = re.sub(r"\[([^\]]+)\]\([^)]+\)", r"\1", body)
sentences = SENTENCE_RE.findall(body)
return [s.strip() for s in sentences if len(s.strip().split()) >= 3]
def burstiness_score(sentences):
"""Compute burstiness (sentence length variance). High = human, low = AI."""
if len(sentences) < 5:
return {"score": 50, "mean": 0, "std": 0, "cv": 0, "note": "too few sentences"}
lengths = [len(s.split()) for s in sentences]
mean = sum(lengths) / len(lengths)
variance = sum((l - mean) ** 2 for l in lengths) / len(lengths)
std = math.sqrt(variance)
cv = std / mean if mean > 0 else 0 # coefficient of variation
# Human writing: CV typically 0.4-0.8+ (high variance)
# AI writing: CV typically 0.15-0.35 (consistently medium)
if cv >= 0.5:
ai_prob = max(0, 30 - (cv - 0.5) * 60)
elif cv >= 0.35:
ai_prob = 30 + (0.5 - cv) * 130
else:
ai_prob = 50 + (0.35 - cv) * 250
ai_prob = max(0, min(100, ai_prob))
return {
"score": round(ai_prob, 1),
"mean_sentence_length": round(mean, 1),
"std_sentence_length": round(std, 1),
"coefficient_of_variation": round(cv, 3),
"note": "low CV = flat sentence lengths (AI-like)" if cv < 0.35 else "healthy variance",
}
def vocabulary_diversity(text):
"""Type-Token Ratio. Low TTR = repetitive vocabulary (AI-like)."""
body = re.sub(r"^---.*?---\s*", "", text, count=1, flags=re.DOTALL)
words = [w.lower() for w in WORD_RE.findall(body) if len(w) > 2]
if len(words) < 50:
return {"score": 50, "ttr": 0, "unique": 0, "total": len(words), "note": "too few words"}
# Use a sliding window TTR for length-independence
window = min(200, len(words))
ttrs = []
for i in range(0, len(words) - window + 1, window // 2):
chunk = words[i : i + window]
ttrs.append(len(set(chunk)) / len(chunk))
avg_ttr = sum(ttrs) / len(ttrs)
# Human: TTR 0.55-0.75+ (varied vocabulary)
# AI: TTR 0.35-0.50 (repetitive patterns)
if avg_ttr >= 0.60:
ai_prob = max(0, 25 - (avg_ttr - 0.60) * 150)
elif avg_ttr >= 0.45:
ai_prob = 25 + (0.60 - avg_ttr) * 330
else:
ai_prob = 75 + (0.45 - avg_ttr) * 250
ai_prob = max(0, min(100, ai_prob))
# Top repeated words (excluding stop words)
stop = {"the", "and", "for", "that", "this", "with", "from", "are", "was", "were", "been",
"have", "has", "had", "not", "but", "can", "will", "your", "you", "they", "their",
"more", "than", "also", "into", "when", "how", "what", "which", "about", "each"}
content_words = [w for w in words if w not in stop]
top_repeated = Counter(content_words).most_common(5)
return {
"score": round(ai_prob, 1),
"avg_ttr": round(avg_ttr, 3),
"unique_words": len(set(words)),
"total_words": len(words),
"top_repeated": [{"word": w, "count": c} for w, c in top_repeated],
"note": "low TTR = repetitive vocabulary (AI-like)" if avg_ttr < 0.45 else "healthy diversity",
}
def phrase_detection(text):
"""Count known AI phrases."""
found = []
text_lower = text.lower()
for i, pattern in enumerate(AI_PHRASE_RES):
matches = pattern.findall(text_lower)
if matches:
found.append({"phrase": AI_PHRASES[i], "count": len(matches)})
total_matches = sum(f["count"] for f in found)
word_count = len(text.split())
density = (total_matches / (word_count / 1000)) if word_count > 0 else 0
# 0-2 per 1K words = normal; 3-5 = suspect; 6+ = likely AI
if density <= 2:
ai_prob = density * 15
elif density <= 5:
ai_prob = 30 + (density - 2) * 20
else:
ai_prob = min(100, 90 + (density - 5) * 5)
return {
"score": round(ai_prob, 1),
"phrases_found": len(found),
"total_matches": total_matches,
"density_per_1k_words": round(density, 1),
"matches": found[:10],
"note": f"{density:.1f} AI phrases per 1K words" if found else "no known AI phrases",
}
def analyze(text):
sentences = extract_sentences(text)
burst = burstiness_score(sentences)
vocab = vocabulary_diversity(text)
phrases = phrase_detection(text)
# Weighted composite: burstiness 35%, vocab 30%, phrases 35%
composite = burst["score"] * 0.35 + vocab["score"] * 0.30 + phrases["score"] * 0.35
composite = round(min(100, max(0, composite)), 1)
if composite <= 20:
verdict = "LIKELY_HUMAN"
elif composite <= 50:
verdict = "MIXED"
else:
verdict = "LIKELY_AI"
return {
"status": "ok",
"composite_score": composite,
"verdict": verdict,
"burstiness": burst,
"vocabulary": vocab,
"phrases": phrases,
"sentences_analyzed": len(sentences),
"recommendations": _recommendations(burst, vocab, phrases),
}
def _recommendations(burst, vocab, phrases):
recs = []
if burst["score"] > 40:
recs.append("Vary sentence lengths: mix short punchy sentences (5-8 words) with longer explanatory ones (20-30 words).")
if vocab["score"] > 40:
recs.append("Diversify vocabulary: replace repeated words with synonyms. Use domain-specific jargon where appropriate.")
if phrases["total_matches"] > 0:
top = phrases["matches"][:3]
recs.append(f"Remove/replace AI phrases: {', '.join(p['phrase'] for p in top)}")
if not recs:
recs.append("Content reads naturally. No humanization needed.")
return recs
def main():
p = argparse.ArgumentParser(
description="Detect AI-generated content via burstiness, vocabulary diversity, and phrase analysis.",
epilog="Score 0-20 = likely human, 21-50 = mixed, 51-100 = likely AI. Run with --demo.",
)
p.add_argument("file", nargs="?", help="Markdown/text file to analyze")
p.add_argument("--json", action="store_true", help="JSON output")
p.add_argument("--demo", action="store_true", help="Run with AI-heavy demo text")
args = p.parse_args()
if args.demo:
text = DEMO_CONTENT
elif args.file:
path = Path(args.file)
if not path.exists():
print(f"[error] {path} not found", file=sys.stderr)
sys.exit(1)
text = path.read_text(encoding="utf-8", errors="replace")
else:
p.print_help()
sys.exit(0)
result = analyze(text)
if args.json:
print(json.dumps(result, indent=2))
return
print(f"AI Content Detection — Composite: {result['composite_score']}/100 ({result['verdict']})")
print()
b = result["burstiness"]
print(f" Burstiness: {b['score']}/100 — CV={b['coefficient_of_variation']} ({b['note']})")
v = result["vocabulary"]
print(f" Vocabulary: {v['score']}/100 — TTR={v['avg_ttr']} ({v['note']})")
ph = result["phrases"]
print(f" AI Phrases: {ph['score']}/100 — {ph['total_matches']} matches, {ph['density_per_1k_words']}/1K words")
if ph["matches"]:
for m in ph["matches"][:5]:
print(f" → \"{m['phrase']}\" (×{m['count']})")
print()
print("Recommendations:")
for r in result["recommendations"]:
print(f" → {r}")
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
readability_scorer.py — Readability metrics for marketing copy
Usage:
python3 readability_scorer.py --file copy.txt
echo "Your text here" | python3 readability_scorer.py
python3 readability_scorer.py # demo mode
python3 readability_scorer.py --json
"""
import argparse
import json
import math
import re
import sys
# ---------------------------------------------------------------------------
# Word lists
# ---------------------------------------------------------------------------
FILLER_WORDS = [
"very", "really", "just", "actually", "basically", "literally",
"honestly", "totally", "absolutely", "definitely", "certainly",
"obviously", "clearly", "quite", "rather", "somewhat", "fairly",
"pretty", "simply", "truly", "genuinely", "essentially",
]
# Simple passive voice detection: "was/were/is/are/been/being + past participle"
PASSIVE_PATTERN = re.compile(
r"\b(was|were|is|are|been|being|be|am)\s+(\w+ed|known|written|built|made|done|seen|given|taken|brought|thought|found|put|set|cut|read|let|hit|hurt|cost|led|felt|kept|left|meant|sent|spent|stood|told|wore|won|beat|lost|broke|chose|drove|flew|froze|grew|hid|rang|rode|rose|ran|sank|sang|spoke|swore|swam|threw|woke|wrote)\b",
re.IGNORECASE,
)
ADVERB_PATTERN = re.compile(r"\b\w+ly\b", re.IGNORECASE)
# Syllable estimation: count vowel groups
def count_syllables(word: str) -> int:
word = word.lower().strip(".,!?;:\"'")
if not word:
return 0
# Silent e
if word.endswith("e") and len(word) > 2:
word = word[:-1]
count = len(re.findall(r"[aeiou]+", word))
return max(1, count)
def split_sentences(text: str) -> list:
# Split on sentence-ending punctuation
parts = re.split(r"(?<=[.!?])\s+", text.strip())
return [p.strip() for p in parts if p.strip()]
def split_words(text: str) -> list:
return re.findall(r"\b[a-zA-Z]+\b", text)
# ---------------------------------------------------------------------------
# Metrics
# ---------------------------------------------------------------------------
def flesch_reading_ease(avg_sentence_len: float, avg_syllables: float) -> float:
"""Flesch Reading Ease formula."""
score = 206.835 - (1.015 * avg_sentence_len) - (84.6 * avg_syllables)
return round(max(0.0, min(100.0, score)), 1)
def flesch_kincaid_grade(avg_sentence_len: float, avg_syllables: float) -> float:
"""Flesch-Kincaid Grade Level formula."""
grade = (0.39 * avg_sentence_len) + (11.8 * avg_syllables) - 15.59
return round(max(0.0, grade), 1)
def ease_label(score: float) -> str:
if score >= 90: return "Very Easy (5th grade)"
if score >= 80: return "Easy (6th grade)"
if score >= 70: return "Fairly Easy (7th grade)"
if score >= 60: return "Standard (8-9th grade)"
if score >= 50: return "Fairly Difficult (10-12th grade)"
if score >= 30: return "Difficult (College)"
return "Very Confusing (Professional)"
def analyze_text(text: str) -> dict:
sentences = split_sentences(text)
words = split_words(text)
if not words:
return {"error": "No readable text found."}
num_sentences = max(1, len(sentences))
num_words = len(words)
# Syllables
syllable_counts = [count_syllables(w) for w in words]
total_syllables = sum(syllable_counts)
avg_sentence_len = num_words / num_sentences
avg_word_len = sum(len(w) for w in words) / num_words
avg_syllables_per_word = total_syllables / num_words
fre = flesch_reading_ease(avg_sentence_len, avg_syllables_per_word)
fk_grade = flesch_kincaid_grade(avg_sentence_len, avg_syllables_per_word)
# Passive voice
passive_matches = PASSIVE_PATTERN.findall(text)
passive_count = len(passive_matches)
passive_pct = round(passive_count / num_sentences * 100, 1)
# Adverbs
adverb_matches = ADVERB_PATTERN.findall(text)
# Filter obvious non-adverbs
non_adverb = {"family", "early", "only", "likely", "nearly", "really",
"daily", "weekly", "monthly", "yearly", "friendly", "lovely",
"lonely", "lively", "elderly", "costly"}
adverbs = [a for a in adverb_matches if a.lower() not in non_adverb]
adverb_density = round(len(adverbs) / num_words * 100, 1)
# Filler words
text_lower = text.lower()
word_tokens_lower = [w.lower() for w in words]
filler_found = {fw: word_tokens_lower.count(fw) for fw in FILLER_WORDS if fw in word_tokens_lower}
filler_total = sum(filler_found.values())
# Scoring:
# FRE already 0-100 (higher = easier = better for marketing copy)
# Target for marketing: 60-80 range
fre_score = fre # use as-is
return {
"stats": {
"word_count": num_words,
"sentence_count": num_sentences,
"avg_sentence_length": round(avg_sentence_len, 1),
"avg_word_length": round(avg_word_len, 1),
"avg_syllables_per_word": round(avg_syllables_per_word, 2),
},
"flesch_reading_ease": {
"score": fre,
"label": ease_label(fre),
"target": "60-80 for most marketing copy",
},
"flesch_kincaid_grade": {
"grade_level": fk_grade,
"note": f"Equivalent to grade {fk_grade} reading level",
},
"passive_voice": {
"count": passive_count,
"percentage": passive_pct,
"target": "<10%",
"pass": passive_pct < 10,
},
"adverb_density": {
"count": len(adverbs),
"percentage": adverb_density,
"examples": list(set(adverbs))[:8],
"target": "<5%",
"pass": adverb_density < 5,
},
"filler_words": {
"total_count": filler_total,
"breakdown": filler_found,
"target": "0-3 per 100 words",
"per_100_words": round(filler_total / num_words * 100, 1),
},
"overall_score": round(fre),
}
# ---------------------------------------------------------------------------
# Demo text
# ---------------------------------------------------------------------------
DEMO_TEXT = """
Marketing copy needs to be clear, direct, and persuasive. When you write for your audience,
you should always think about what they actually want to hear. Really good copy is basically
about solving problems. It is very important to avoid using overly complicated language that
might confuse the reader.
The best headlines are written by experts who truly understand their customers. A strong
call-to-action is absolutely essential for any landing page. You need to make sure that
every single word is earning its place on the page.
Studies show that shorter sentences improve comprehension. The average reader processes
information faster when sentences contain fewer than 20 words. This is genuinely proven
by research. Passive voice constructions are often used by writers who want to sound
authoritative, but they can actually make copy feel distant and unclear.
Focus on benefits, not features. Tell the reader what they will gain. Use numbers when
you can — "save 3 hours per week" beats "save time" every single time. Specificity
builds trust. Vague promises are ignored.
"""
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(
description="Readability scorer for marketing copy — Flesch, passive voice, filler words."
)
parser.add_argument("--file", help="Path to text file")
parser.add_argument("--json", action="store_true", help="Output as JSON")
args = parser.parse_args()
if args.file:
with open(args.file, "r", encoding="utf-8", errors="replace") as f:
text = f.read()
elif not sys.stdin.isatty():
text = sys.stdin.read()
if not text.strip():
text = DEMO_TEXT
if not args.json:
print("No input provided — running in demo mode.\n")
else:
text = DEMO_TEXT
if not args.json:
print("No input provided — running in demo mode.\n")
result = analyze_text(text)
if "error" in result:
print(f"Error: {result['error']}", file=sys.stderr)
sys.exit(1)
if args.json:
print(json.dumps(result, indent=2))
return
fre = result["flesch_reading_ease"]
fk = result["flesch_kincaid_grade"]
stats = result["stats"]
passive = result["passive_voice"]
adverbs = result["adverb_density"]
fillers = result["filler_words"]
score = result["overall_score"]
PASS = "✅"
FAIL = "❌"
print("=" * 62)
print(f" READABILITY REPORT Flesch Score: {fre['score']}/100")
print("=" * 62)
print(f" {fre['label']}")
print(f" Target: {fre['target']}")
print()
print(f" 📊 Stats")
print(f" Words: {stats['word_count']}")
print(f" Sentences: {stats['sentence_count']}")
print(f" Avg sentence length:{stats['avg_sentence_length']} words")
print(f" Avg word length: {stats['avg_word_length']} chars")
print(f" Syllables/word: {stats['avg_syllables_per_word']}")
print()
print(f" 📐 Flesch-Kincaid Grade Level: {fk['grade_level']}")
print(f" {fk['note']}")
print()
pv_icon = PASS if passive["pass"] else FAIL
print(f" {pv_icon} Passive Voice: {passive['count']} instances ({passive['percentage']}%)")
print(f" Target: {passive['target']}")
av_icon = PASS if adverbs["pass"] else FAIL
print(f" {av_icon} Adverb Density: {adverbs['count']} adverbs ({adverbs['percentage']}%)")
if adverbs["examples"]:
print(f" Examples: {', '.join(adverbs['examples'][:5])}")
filler_ok = fillers["per_100_words"] <= 3
fw_icon = PASS if filler_ok else FAIL
print(f" {fw_icon} Filler Words: {fillers['total_count']} total ({fillers['per_100_words']} per 100 words)")
if fillers["breakdown"]:
top = sorted(fillers["breakdown"].items(), key=lambda x: -x[1])[:5]
print(f" Top: {', '.join(f'{w}({c})' for w,c in top)}")
print()
print("=" * 62)
score_bar_len = round(score / 10)
bar = "█" * score_bar_len + "░" * (10 - score_bar_len)
print(f" Readability Score: [{bar}] {score}/100")
print("=" * 62)
if __name__ == "__main__":
main()
Related skills
FAQ
What metrics does copy-editing use to detect AI text?
copy-editing uses ai_content_detector.py with three methods, starting with burstiness—the coefficient of variation of sentence lengths. CV 0.50+ indicates natural human rhythm (0–30% AI probability); CV 0.35–0.49 suggests mixed or flat AI-like pacing.
What content types does copy-editing humanize?
copy-editing humanizes landing pages, marketing emails, and documentation flagged as AI-flat. The skill targets varied sentence lengths—short 4–8 word punches and longer 20–35 word explanations—instead of uniform 12–20 word AI rhythm.
Is Copy Editing safe to install?
skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.