
Copy Editing
- 100 installs
- 451 repo stars
- Updated July 21, 2026
- borghei/claude-skills
Copy Editing is a Claude skill that improves marketing copy through seven focused editorial sweeps covering clarity, voice, proof, specificity, emotion, and friction removal.
About
Copy Editing improves marketing copy through focused editorial passes covering clarity, voice consistency, benefit framing, proof validation, specificity, emotional impact, and friction removal. A writer or editor uses it to review drafts, polish content, proofread, or fact-check. Its core method is the Seven Sweeps Framework plus editorial checklists and style-consistency standards.
- Seven Sweeps Framework: clarity, voice, benefit framing, proof, specificity, emotion, and friction removal
- Includes editorial checklists, common-problem diagnosis, and style-consistency standards
- Fact-checking protocol to validate claims and proof
Copy Editing by the numbers
- 100 all-time installs (skills.sh)
- Ranked #1,148 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
copy-editing capabilities & compatibility
- Capabilities
- copy editing · proofreading · fact checking
- Use cases
- copywriting · marketing
- Pricing
- Free
What copy-editing says it does
Systematic copy improvement through focused editorial passes that enhance clarity, voice, proof, and conversion impact.
Edit through seven sequential passes. Each focuses on one dimension.
Sentence over 30 words | Split into two sentences
npx skills add https://github.com/borghei/claude-skills --skill copy-editingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 100 |
|---|---|
| repo stars | ★ 451 |
| Last updated | July 21, 2026 |
| Repository | borghei/claude-skills ↗ |
What it does
Edit and polish marketing copy through seven focused sweeps covering clarity, proof, and conversion.
Who is it for?
Editing, proofreading, and polishing marketing copy and drafts
Skip if: Writing copy from scratch or code review
When should I use this skill?
Editing marketing copy, reviewing drafts, polishing content, or proofreading
What you get
Produces clearer, more consistent, higher-converting copy validated against an editorial checklist.
- edited copy
- editorial checklist review
- voice-consistency fixes
By the numbers
- Seven Sweeps editorial framework
- Flags sentences over 30 words for splitting
Files
Copy Editing
Systematic copy improvement through focused editorial passes that enhance clarity, voice, proof, and conversion impact.
---
Table of Contents
- Keywords
- Quick Start
- The Seven Sweeps Framework
- Quick-Pass Editing Guide
- Common Copy Problems and Fixes
- Style Consistency Standards
- Fact-Checking Protocol
- Editorial Checklist
- Best Practices
- Integration Points
---
Keywords
copy editing, editorial review, copy feedback, proofreading, content polishing, copy sweep, editorial standards, style consistency, grammar check, fact-checking, clarity editing, voice consistency, benefit framing, proof validation, specificity, conversion copy editing, marketing copy review, copy quality
---
Quick Start
Full Copy Review (Seven Sweeps)
1. Read through once without editing to understand the whole piece 2. Sweep 1 — Clarity: flag confusing sentences, unclear references, jargon 3. Sweep 2 — Voice and Tone: flag shifts in formality, personality inconsistencies 4. Sweep 3 — So What: flag features without benefits, claims without consequences 5. Sweep 4 — Prove It: flag unsubstantiated claims, missing social proof 6. Sweep 5 — Specificity: flag vague language, round numbers, generic statements 7. Sweep 6 — Heightened Emotion: strengthen pain points, aspirations, urgency 8. Sweep 7 — Zero Risk: remove barriers near CTAs, add trust signals
Quick Copy Pass
1. Cut filler words (very, really, just, actually, basically) 2. Replace weak verbs (utilize > use, facilitate > help, leverage > use) 3. Fix passive voice (reports are generated > we generate reports) 4. Verify one idea per sentence, one topic per paragraph 5. Check CTA for action orientation
---
The Seven Sweeps Framework
Edit through seven sequential passes. Each focuses on one dimension. After each sweep, verify previous sweeps are not compromised.
Sweep 1: Clarity
Focus: Can the reader understand what you are saying on the first read?
What to check:
- Sentences trying to say too much (split them)
- Unclear pronoun references ("it" — what is "it"?)
- Jargon or insider language without explanation
- Ambiguous statements that could be read two ways
- Missing context that assumes reader knowledge
Clarity killers to fix:
| Problem | Fix |
|---|---|
| Sentence over 30 words | Split into two sentences |
| Abstract language | Replace with concrete example |
| Buried main point | Move to beginning of paragraph |
| Three-clause sentence | Simplify to one or two clauses |
| Undefined acronym | Spell out on first use |
After this sweep: Confirm the "Rule of One" (one idea per section) and "You Rule" (copy speaks to the reader as "you") are intact.
Sweep 2: Voice and Tone
Focus: Does the copy sound consistent throughout?
What to check:
- Shifts between formal and casual language
- Inconsistent brand personality (joking in one paragraph, corporate in the next)
- Jarring mood changes without transition
- Word choices that do not match the established voice
- Mixing "we" and "the company" references
Voice consistency indicators:
| Consistent | Inconsistent |
|---|---|
| Same level of contractions throughout | Contractions in some sections, full forms in others |
| Humor style maintained | Random joke in otherwise serious copy |
| Same sentence structure patterns | Short punchy intro, corporate middle, casual close |
| Consistent use of "you" | Switching between "you," "users," "customers," "one" |
After this sweep: Return to Sweep 1 to ensure voice edits did not introduce confusion.
Sweep 3: So What
Focus: Does every claim answer "why should I care?"
The So What test: For every statement, ask "So what?" If the copy does not answer with a deeper benefit, it needs work.
| Before (features only) | After (feature + benefit) |
|---|---|
| "Our platform uses AI-powered analytics" | "Our AI analytics surface insights you would miss manually — so you make better decisions in half the time" |
| "SOC 2 Type II certified" | "SOC 2 certified — your security team approves us in days, not months" |
| "Real-time dashboard" | "See exactly what is happening right now, not what happened last week" |
After this sweep: Return to Sweeps 2 and 1.
Sweep 4: Prove It
Focus: Is every claim backed with evidence?
Types of proof to verify:
| Proof Type | Strength | Example |
|---|---|---|
| Named testimonial | Strong | "Sarah Chen, VP Marketing at Stripe: 'Reduced our setup time by 60%'" |
| Specific statistic | Strong | "2,847 teams use [Product] daily" |
| Case study reference | Strong | "See how Linear reduced churn by 23% in 90 days" |
| Third-party validation | Strong | "Named a Leader in Gartner's Magic Quadrant 2025" |
| Customer logos | Medium | Recognizable brand logos with permission |
| Generic claim | Weak — flag it | "Customers love us," "Industry-leading" |
Common proof gaps to flag:
- "Trusted by thousands" (which thousands? give a number)
- "Industry-leading" (according to whom? cite the source)
- "Best-in-class" (by what measure?)
- "Customers love us" (show them saying it)
- Results claims without timeframe or specifics
After this sweep: Return to Sweeps 3, 2, and 1.
Sweep 5: Specificity
Focus: Is the copy concrete enough to be compelling?
Specificity upgrades:
| Vague | Specific |
|---|---|
| "Save time" | "Save 4 hours every week" |
| "Many customers" | "2,847 teams" |
| "Fast results" | "Results in 14 days" |
| "Improve your workflow" | "Cut reporting time from 4 hours to 15 minutes" |
| "Great support" | "Average response time: 2 hours" |
| "Easy to use" | "Set up in 10 minutes, no code required" |
| "Affordable" | "Starting at $29/month" |
| "Scalable" | "Handles 10,000 to 10 million records without slowdown" |
Rule: If a claim cannot be made specific, it is probably filler. Cut it or replace it with something verifiable.
After this sweep: Return to Sweeps 4, 3, 2, and 1.
Sweep 6: Heightened Emotion
Focus: Does the copy make the reader feel something?
Emotional dimensions to check:
| Emotion | Where to Use | Technique |
|---|---|---|
| Pain/frustration | Problem section | Paint the "before" state vividly |
| Relief | Solution section | Show the contrast with current pain |
| Fear of missing out | Social proof | "Teams like yours already use..." |
| Pride | Aspiration section | "Be the team that..." |
| Confidence | CTA area | "Join 2,847 teams who already..." |
| Urgency | Near CTA | Only if genuine (real deadline, limited spots) |
Emotion techniques:
- Paint the "before" state with sensory detail
- Use micro-stories (1-2 sentences) from customer scenarios
- Ask questions that prompt self-reflection
- Reference shared experiences the audience recognizes
After this sweep: Return to Sweeps 5, 4, 3, 2, and 1.
Sweep 7: Zero Risk
Focus: Have we removed every barrier to action?
Friction checklist near CTAs:
- [ ] What happens after clicking is clear (not a mystery)
- [ ] Objections addressed within 2 scrolls of the CTA
- [ ] Trust signals visible (guarantee, certifications, customer count)
- [ ] Next steps are specific ("Start your 14-day free trial" not "Get started")
- [ ] Risk reversals stated explicitly (money-back, no CC, cancel anytime)
- [ ] Privacy concerns addressed if form collects data
After this sweep: Return through all previous sweeps one final time.
---
Quick-Pass Editing Guide
Word-Level Cuts
Always cut: very, really, extremely, incredibly, quite, rather, somewhat, just, actually, basically, essentially, literally (unless literal), in order to (use "to"), the fact that, it should be noted that, it is important to
Always replace:
| Weak | Strong |
|---|---|
| Utilize | Use |
| Implement | Set up, build, create |
| Leverage | Use, apply |
| Facilitate | Help, enable |
| Innovative | New, original, first-of-its-kind |
| Robust | Strong, thorough, [be specific] |
| Seamless | Smooth, easy, [be specific] |
| Cutting-edge | Modern, latest, [be specific] |
| Synergy | [Delete or be specific about the collaboration] |
| Paradigm | [Delete or say what actually changed] |
Sentence-Level Checks
- One idea per sentence
- Vary sentence length (mix 8-word and 20-word sentences)
- Front-load important information (do not bury the point)
- Maximum 3 conjunctions per sentence
- Active voice default (flip passive constructions)
Paragraph-Level Checks
- One topic per paragraph
- 2-4 sentences maximum for web copy
- Strong opening sentence that states the paragraph's point
- Logical flow between paragraphs
- White space for scannability
---
Common Copy Problems and Fixes
| Problem | Symptom | Fix |
|---|---|---|
| Wall of features | List of what it does, no why | Add "which means..." after each feature |
| Corporate speak | "Leverage synergies to optimize outcomes" | Ask "How would a human say this?" |
| Weak opening | Starts with company history or vague statement | Lead with reader's problem or desired outcome |
| Buried CTA | Ask comes after too much buildup | Make CTA obvious, early, and repeated |
| No proof | "Customers love us" with no evidence | Add specific testimonials, numbers, case references |
| Generic claims | "We help businesses grow" | Specify who, how, and by how much |
| Mixed audiences | Tries to speak to everyone | Pick one audience per page/section |
| Feature overload | Every capability listed | Focus on 3-5 benefits that matter most |
| Passive voice | "Reports are generated by the system" | "The system generates reports" |
| Weasel words | "Up to 50% improvement" | State the median or typical result with context |
---
Style Consistency Standards
Checklist for Style Consistency
- [ ] Oxford comma: used consistently (or consistently omitted)
- [ ] Contractions: consistent usage throughout
- [ ] Heading case: consistent (sentence case or title case, not mixed)
- [ ] Number style: consistent (spell out 1-9, numerals for 10+, or chosen standard)
- [ ] Date format: consistent (March 9, 2026 or 2026-03-09, not mixed)
- [ ] Brand name: capitalized and formatted consistently
- [ ] Product/feature names: capitalized consistently per brand standards
- [ ] Bulleted lists: consistent punctuation (periods or no periods)
- [ ] Acronyms: spelled out on first use in each document
- [ ] Em dashes, en dashes, hyphens: used correctly and consistently
- [ ] Quotation marks: consistent style (straight or curly)
Common Style Conflicts
| Decision | Option A | Option B | How to Decide |
|---|---|---|---|
| Oxford comma | Yes | No | Pick one, document it, enforce it |
| Heading capitalization | Sentence case | Title Case | Sentence case is modern standard |
| "Login" vs "Log in" | One word (noun) | Two words (verb) | "Log in" as verb, "Login" as noun/adjective |
| "Setup" vs "Set up" | One word (noun) | Two words (verb) | Same pattern as login |
| Ampersand vs "and" | & | and | "and" in prose, "&" only in headings if brand standard |
---
Fact-Checking Protocol
What to Verify
- [ ] All statistics have a named source and year
- [ ] Customer testimonials are attributed to real, named individuals
- [ ] Customer logos are used with permission
- [ ] Competitive claims are accurate and current
- [ ] Product capabilities described are actually available (not roadmap items)
- [ ] Pricing is current and matches the pricing page
- [ ] Certifications and compliance claims are active (SOC 2, GDPR, etc.)
- [ ] Awards and recognition are current year or specified year
- [ ] Integration claims list actual integrations, not aspirational ones
- [ ] Uptime/SLA claims match the actual SLA
Red Flags to Investigate
- Round numbers without source (sounds made up)
- Superlatives without qualification ("fastest," "best," "only")
- Claims that contradict other pages on the same site
- Screenshots from an older version of the product
- Competitor comparisons without date (may be outdated)
---
Editorial Checklist
Pre-Edit
- [ ] Understand the goal of this copy
- [ ] Know the target audience
- [ ] Identify the desired action
- [ ] Read through once without editing
During Edit (Seven Sweeps Summary)
- [ ] Sweep 1: Every sentence is immediately understandable
- [ ] Sweep 2: Voice is consistent throughout
- [ ] Sweep 3: Every feature connects to a benefit
- [ ] Sweep 4: Claims are substantiated with evidence
- [ ] Sweep 5: Vague words replaced with specifics
- [ ] Sweep 6: Copy evokes appropriate emotion
- [ ] Sweep 7: Barriers to action removed near CTAs
Post-Edit
- [ ] No typos or grammatical errors
- [ ] Consistent formatting throughout
- [ ] Core message preserved through all edits
- [ ] All links functional (if applicable)
- [ ] Consistent style applied (see style checklist)
---
Best Practices
1. Edit in passes, not all at once — Trying to fix everything in one read misses issues. Each sweep catches what the others miss.
2. Preserve the author's voice — Good copy editing enhances; it does not replace. Maintain the original voice while improving clarity and impact.
3. Every edit needs a reason — Never change a word without explaining the principle. "Changed because it is clearer" or "Replaced because the original was vague."
4. Prioritize by conversion impact — Fix the CTA before fixing a comma. Fix the headline before fixing paragraph 12.
5. Read aloud — Voice and rhythm problems become obvious when read aloud. If it sounds wrong spoken, it reads wrong too.
6. Flag what you cannot fix — If a claim needs proof the author must provide, flag it clearly. You can improve phrasing but you cannot invent evidence.
7. Re-check previous sweeps — Each sweep can introduce issues caught by earlier sweeps. Always go back.
8. Cut first, add second — Most marketing copy is 20-30% too long. Cut the fat before adding new content.
9. Get context before editing — A copy edit without knowing the audience, goal, and voice standard produces misaligned feedback.
10. Track recurring issues — If the same problems appear across multiple pieces, the issue is systemic. Flag it as a process improvement, not just an edit.
---
Integration Points
- Copywriting — Use for writing new copy from scratch. Copy Editing handles reviewing and improving existing copy.
- Content Humanizer — Use when AI-generated copy needs humanization before editorial review.
- Content Production — Use Copy Editing as part of the production pipeline between drafting and publishing.
- Brand Guidelines — Reference brand voice and style standards during the Voice and Tone sweep.
- Marketing Psychology — Apply psychological principles during the Heightened Emotion sweep.
- Content Strategy — Use when the problem is what to say, not how to say it.
---
Troubleshooting
| Problem | Likely Cause | Fix |
|---|---|---|
| Same issues appear across multiple pieces from the same writer | Systemic writing habit, not a one-off error | Create a writer-specific checklist of recurring issues; address in style guide or training, not just per-piece edits |
| Copy loses its original voice after editing | Editor over-corrected; replaced author voice with editor's style | Preserve author voice — enhance clarity and impact without rewriting personality. Read original aloud before editing |
| Edits introduce new inconsistencies | Previous sweeps not re-checked after later sweeps modified content | Always re-run earlier sweeps after making changes — Sweep 7 edits can break Sweep 1 clarity |
| CTA buried or ineffective despite multiple edit passes | CTA was not the focus of any sweep — copy editing focused on prose quality | Prioritize CTA area first (Sweep 7: Zero Risk) before polishing earlier sections |
| Fact-checking reveals unverifiable claims | Writer invented statistics or used outdated data | Flag and return to writer — editor cannot invent evidence. Document all unverifiable claims explicitly |
| Style inconsistencies between sections | Multiple writers contributed or copy was assembled from different drafts | Run style consistency checklist end-to-end; standardize contractions, heading case, number format, and punctuation |
---
Success Criteria
- Seven Sweeps completion: All 7 sweeps completed per piece with previous sweeps re-verified after each pass
- Error rate: Zero grammatical errors, typos, or broken links in published copy
- Style consistency: 100% adherence to documented style guide (Oxford comma, heading case, number format, etc.)
- Fact verification: All statistics have named source and year; all claims are verifiable or labeled as opinion
- CTA effectiveness: Every piece has a clear, specific CTA visible within 2 scrolls of the content end
- Clarity score: Every sentence understandable on first read — zero ambiguous pronoun references or multi-clause confusion
- Edit turnaround: 48-hour maximum turnaround on editorial review with clear change documentation
---
Scope & Limitations
In scope:
- Seven Sweeps editorial framework (clarity, voice, so-what, proof, specificity, emotion, zero-risk)
- Quick-pass editing (word-level, sentence-level, paragraph-level)
- Style consistency auditing and enforcement
- Fact-checking protocol for claims, statistics, and competitive references
- Pre-edit and post-edit checklists
- Common copy problem diagnosis and fixing
Out of scope:
- Writing new copy from scratch (use Copywriting skill)
- AI content detection and humanization (use Content Humanizer)
- SEO optimization passes (use Content Production optimization pipeline)
- Content strategy or topic selection (use Content Strategy)
- Visual design or layout feedback
- Legal review of marketing claims
Known limitations:
- Cannot verify internal company claims (product capabilities, uptime SLAs) without access to product documentation
- Fact-checking external claims requires access to original sources — may need writer input
- Style consistency requires an existing style guide; without one, editor must make judgment calls
- Emotional impact (Sweep 6) is subjective and varies by audience — use target audience context
- Multi-language copy editing requires native-level proficiency in each language
---
Scripts
# Score content readability with detailed metrics
python scripts/readability_scorer.py article.md --json
# Detect AI patterns that need humanization before editing
python scripts/ai_pattern_detector.py article.md --verbose
# Check style consistency across multiple documents
python scripts/style_checker.py --files doc1.md doc2.md doc3.md --json#!/usr/bin/env python3
"""
AI Pattern Detector for Copy Editing
Detects AI-generated content markers relevant to editorial review.
Identifies filler words, hedging, structural uniformity, and provides
a pre-edit assessment for copy editors to know what they are working with.
Usage:
python ai_pattern_detector.py article.md
python ai_pattern_detector.py article.md --json
python ai_pattern_detector.py article.md --verbose
"""
import argparse
import json
import math
import re
import sys
from pathlib import Path
from collections import Counter
FILLER_CRITICAL = [
'delve', 'landscape', 'crucial', 'vital', 'pivotal', 'leverage',
'robust', 'comprehensive', 'holistic', 'foster', 'facilitate',
'utilize', 'furthermore', 'moreover', 'navigate', 'embark',
'tapestry', 'multifaceted', 'underscore',
]
FILLER_MEDIUM = [
'streamline', 'optimize', 'innovative', 'cutting-edge', 'game-changer',
'paradigm', 'synergy', 'ecosystem', 'empower', 'transformative',
'seamless', 'elevate', 'spearhead', 'groundbreaking',
]
HEDGING = [
"it's important to note", "it is important to note",
"it's worth mentioning", "it is worth mentioning",
"one might argue", "in many cases",
"it goes without saying", "needless to say",
]
GENERIC_OPENERS = [
r'^in today\'?s (digital|modern|fast-paced|competitive|evolving)',
r'^in the (rapidly|ever|constantly) (evolving|changing)',
r'^in an? (increasingly|highly|rapidly)',
]
def detect_patterns(text):
"""Detect all AI patterns."""
wc = len(text.split())
findings = []
# Filler words
for word in FILLER_CRITICAL:
count = len(re.findall(r'\b' + word + r'\b', text, re.IGNORECASE))
if count:
findings.append({
"category": "filler",
"severity": "Critical",
"item": word,
"count": count,
"fix": f"Replace '{word}' — see content humanizer replacement guide",
})
for word in FILLER_MEDIUM:
count = len(re.findall(r'\b' + word + r'\b', text, re.IGNORECASE))
if count:
findings.append({
"category": "filler",
"severity": "Medium",
"item": word,
"count": count,
"fix": f"Consider replacing '{word}' with plain-language alternative",
})
# Hedging
for phrase in HEDGING:
count = len(re.findall(re.escape(phrase), text, re.IGNORECASE))
if count:
findings.append({
"category": "hedging",
"severity": "Critical",
"item": phrase,
"count": count,
"fix": f"Delete '{phrase}' and start with the actual point",
})
# Generic openers
for pattern in GENERIC_OPENERS:
if re.search(pattern, text, re.IGNORECASE | re.MULTILINE):
findings.append({
"category": "generic_opener",
"severity": "High",
"item": "Generic AI opener detected",
"count": 1,
"fix": "Rewrite opening — lead with the problem or the answer",
})
# Em-dash overuse
em_count = text.count('—') + text.count(' -- ')
per_500 = round(em_count / max(wc, 1) * 500, 1)
if per_500 > 2:
findings.append({
"category": "em_dash",
"severity": "Medium",
"item": f"Em-dash overuse: {em_count} ({per_500} per 500 words)",
"count": em_count,
"fix": "Replace some em-dashes with periods, commas, or parentheses",
})
# Sentence uniformity
sents = [s.strip() for s in re.split(r'[.!?]+', text) if s.strip() and len(s.split()) >= 3]
if len(sents) >= 10:
lengths = [len(s.split()) for s in sents]
std = math.sqrt(sum((l - sum(lengths)/len(lengths)) ** 2 for l in lengths) / len(lengths))
if std < 3.5:
findings.append({
"category": "uniformity",
"severity": "High",
"item": f"Sentence length uniformity: std dev {round(std, 1)}",
"count": 1,
"fix": "Vary sentence length — add short sentences, fragments, and questions",
})
total = sum(f.get("count", 1) for f in findings)
per_500_total = round(total / max(wc, 1) * 500, 1)
if per_500_total < 3:
assessment = "Light editing — minor AI patterns"
elif per_500_total < 7:
assessment = "Moderate editing — noticeable AI patterns throughout"
else:
assessment = "Heavy editing or rewrite — extensive AI patterns"
return {
"word_count": wc,
"total_tells": total,
"per_500_words": per_500_total,
"assessment": assessment,
"findings": findings,
}
def main():
parser = argparse.ArgumentParser(description="Detect AI patterns for copy editing")
parser.add_argument("file", help="Content file")
parser.add_argument("--json", action="store_true")
parser.add_argument("--verbose", action="store_true")
args = parser.parse_args()
fp = Path(args.file)
if not fp.exists():
print(f"Error: {fp} not found", file=sys.stderr)
sys.exit(1)
text = fp.read_text(encoding="utf-8", errors="replace")
result = detect_patterns(text)
result["file"] = str(fp)
if args.json:
print(json.dumps(result, indent=2))
else:
print(f"\n{'='*55}")
print(f" AI PATTERN SCAN — Pre-Edit Assessment")
print(f"{'='*55}")
print(f" File: {fp} | Words: {result['word_count']}")
print(f" AI tells: {result['total_tells']} ({result['per_500_words']} per 500 words)")
print(f" Assessment: {result['assessment']}")
if result["findings"]:
# Group by category
categories = {}
for f in result["findings"]:
cat = f["category"]
if cat not in categories:
categories[cat] = []
categories[cat].append(f)
for cat, items in categories.items():
total = sum(i.get("count", 1) for i in items)
print(f"\n {cat.upper()} ({total} instances):")
for item in items:
sev = item["severity"]
if args.verbose:
print(f" [{sev}] {item['item']} (x{item['count']})")
print(f" Fix: {item['fix']}")
else:
print(f" [{sev}] {item['item']} (x{item['count']})")
else:
print(f"\n No AI patterns detected — content reads as human-written.")
print()
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Copy Editing Readability Scorer
Scores content for editorial readability including Flesch Reading Ease,
sentence complexity, paragraph structure, passive voice rate, filler
word density, and web-optimized formatting checks.
Usage:
python readability_scorer.py article.md
python readability_scorer.py article.md --json
python readability_scorer.py article.md --verbose
"""
import argparse
import json
import math
import re
import sys
from pathlib import Path
def strip_formatting(text):
"""Remove markdown formatting."""
t = re.sub(r'^#{1,6}\s+', '', text, flags=re.MULTILINE)
t = re.sub(r'\*\*([^*]+)\*\*', r'\1', t)
t = re.sub(r'\*([^*]+)\*', r'\1', t)
t = re.sub(r'```.*?```', '', t, flags=re.DOTALL)
t = re.sub(r'`[^`]+`', '', t)
t = re.sub(r'\[([^\]]+)\]\([^)]+\)', r'\1', t)
t = re.sub(r'^\|.*\|$', '', t, flags=re.MULTILINE)
return t.strip()
def count_syllables(word):
"""Estimate syllables."""
w = word.lower().strip('.,!?;:()[]"\'-')
if len(w) <= 3:
return 1
count = 0
prev = False
for c in w:
v = c in 'aeiouy'
if v and not prev:
count += 1
prev = v
if w.endswith('e') and count > 1:
count -= 1
return max(count, 1)
def sentences(text):
"""Get sentences."""
return [s.strip() for s in re.split(r'(?<=[.!?])\s+', text) if s.strip() and len(s.split()) >= 3]
def flesch(words, sents, syls):
"""Flesch Reading Ease."""
if words == 0 or sents == 0:
return 0
return round(206.835 - 1.015 * (words / sents) - 84.6 * (syls / words), 1)
FILLER_WORDS = [
'very', 'really', 'extremely', 'incredibly', 'quite', 'rather',
'somewhat', 'just', 'actually', 'basically', 'essentially',
'literally', 'simply', 'totally', 'absolutely', 'definitely',
]
WEAK_VERBS = [
'utilize', 'implement', 'leverage', 'facilitate', 'optimize',
'streamline', 'enhance', 'maximize', 'synergize',
]
def analyze(text):
"""Full readability analysis."""
clean = strip_formatting(text)
words = clean.split()
wc = len(words)
sents = sentences(clean)
sc = len(sents)
syls = sum(count_syllables(w) for w in words)
fre = flesch(wc, sc, syls)
fkg = round(0.39 * (wc / max(sc, 1)) + 11.8 * (syls / max(wc, 1)) - 15.59, 1) if wc > 0 else 0
# Sentence analysis
lengths = [len(s.split()) for s in sents] if sents else [0]
avg_sent = round(sum(lengths) / max(len(lengths), 1), 1)
std_sent = round(math.sqrt(sum((l - avg_sent) ** 2 for l in lengths) / max(len(lengths), 1)), 1) if len(lengths) > 1 else 0
long_sents = sum(1 for l in lengths if l > 30)
# Passive voice
passive = sum(1 for s in sents if re.search(r'\b(is|are|was|were|been|being)\s+\w+(ed|en)\b', s, re.I))
passive_rate = round(passive / max(sc, 1) * 100, 1)
# Filler words
filler_count = sum(len(re.findall(r'\b' + w + r'\b', clean, re.I)) for w in FILLER_WORDS)
filler_per_1000 = round(filler_count / max(wc, 1) * 1000, 1)
# Weak verbs
weak_count = sum(len(re.findall(r'\b' + w + r'\b', clean, re.I)) for w in WEAK_VERBS)
# Paragraphs
paras = [p.strip() for p in re.split(r'\n\s*\n', clean) if p.strip()]
para_sents = [len(sentences(p)) for p in paras]
long_paras = sum(1 for c in para_sents if c > 5)
# Subheading frequency
headings = len(re.findall(r'^#{1,6}\s+', text, re.MULTILINE))
words_per_heading = round(wc / max(headings, 1))
# Score
score = 0
checks = {}
# Flesch (25 pts)
if 55 <= fre <= 75:
score += 25
checks["flesch"] = ("PASS", f"Flesch: {fre} (web-friendly range)")
elif 45 <= fre <= 80:
score += 15
checks["flesch"] = ("PASS", f"Flesch: {fre} (acceptable)")
else:
score += 5
checks["flesch"] = ("FAIL", f"Flesch: {fre} (target 60-70)")
# Sentence length (20 pts)
if 12 <= avg_sent <= 22:
score += 20
checks["sentence_length"] = ("PASS", f"Avg {avg_sent} words/sentence")
else:
score += 8
checks["sentence_length"] = ("FAIL", f"Avg {avg_sent} words/sentence (target 15-20)")
# Sentence variety (15 pts)
if std_sent >= 5:
score += 15
checks["sentence_variety"] = ("PASS", f"Good variety (std dev {std_sent})")
elif std_sent >= 3:
score += 10
checks["sentence_variety"] = ("PASS", f"Moderate variety (std dev {std_sent})")
else:
score += 3
checks["sentence_variety"] = ("FAIL", f"Low variety (std dev {std_sent}, target >5)")
# Passive voice (15 pts)
if passive_rate < 10:
score += 15
checks["passive_voice"] = ("PASS", f"Passive: {passive_rate}%")
elif passive_rate < 20:
score += 8
checks["passive_voice"] = ("FAIL", f"Passive: {passive_rate}% (target <10%)")
else:
score += 3
checks["passive_voice"] = ("FAIL", f"Passive: {passive_rate}% (excessive)")
# Filler words (10 pts)
if filler_per_1000 < 5:
score += 10
checks["filler_words"] = ("PASS", f"{filler_count} filler words ({filler_per_1000}/1000)")
else:
score += 3
checks["filler_words"] = ("FAIL", f"{filler_count} filler words ({filler_per_1000}/1000)")
# Paragraph length (10 pts)
if long_paras == 0:
score += 10
checks["paragraph_length"] = ("PASS", "All paragraphs under 5 sentences")
else:
score += 3
checks["paragraph_length"] = ("FAIL", f"{long_paras} paragraphs exceed 5 sentences")
# Formatting (5 pts)
if words_per_heading <= 350:
score += 5
checks["heading_frequency"] = ("PASS", f"Heading every ~{words_per_heading} words")
else:
checks["heading_frequency"] = ("FAIL", f"Heading every ~{words_per_heading} words (target <350)")
return {
"word_count": wc,
"sentence_count": sc,
"score": min(score, 100),
"flesch_reading_ease": fre,
"flesch_kincaid_grade": fkg,
"avg_sentence_length": avg_sent,
"sentence_std_dev": std_sent,
"long_sentences": long_sents,
"passive_voice_rate": passive_rate,
"filler_word_count": filler_count,
"weak_verb_count": weak_count,
"long_paragraphs": long_paras,
"words_per_heading": words_per_heading,
"checks": checks,
}
def main():
parser = argparse.ArgumentParser(description="Score readability for copy editing")
parser.add_argument("file", help="Content file")
parser.add_argument("--json", action="store_true")
parser.add_argument("--verbose", action="store_true")
args = parser.parse_args()
fp = Path(args.file)
if not fp.exists():
print(f"Error: {fp} not found", file=sys.stderr)
sys.exit(1)
text = fp.read_text(encoding="utf-8", errors="replace")
result = analyze(text)
result["file"] = str(fp)
grade = "A" if result["score"] >= 80 else "B" if result["score"] >= 65 else "C" if result["score"] >= 50 else "D"
if args.json:
# Convert checks tuples to dicts for JSON
result["checks"] = {k: {"status": v[0], "detail": v[1]} for k, v in result["checks"].items()}
print(json.dumps(result, indent=2))
else:
print(f"\n{'='*55}")
print(f" READABILITY: {result['score']}/100 (Grade: {grade})")
print(f"{'='*55}")
print(f" Words: {result['word_count']} | Sentences: {result['sentence_count']}")
print(f" Flesch: {result['flesch_reading_ease']} | Grade: {result['flesch_kincaid_grade']}")
for key, (status, detail) in result["checks"].items():
print(f" [{status}] {detail}")
if args.verbose:
print(f"\n Details:")
print(f" Long sentences (>30 words): {result['long_sentences']}")
print(f" Weak verbs: {result['weak_verb_count']}")
print(f" Filler words: {result['filler_word_count']}")
print()
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Style Consistency Checker
Checks content for style consistency including Oxford comma usage,
heading capitalization, contraction consistency, number formatting,
punctuation patterns, and brand name formatting across one or more files.
Usage:
python style_checker.py --file article.md
python style_checker.py --files doc1.md doc2.md doc3.md --json
python style_checker.py --file article.md --verbose
"""
import argparse
import json
import re
import sys
from pathlib import Path
def check_oxford_comma(text):
"""Check Oxford comma consistency."""
# Pattern: X, Y, and Z (Oxford) vs X, Y and Z (no Oxford)
with_oxford = len(re.findall(r'\w+,\s+\w+,\s+and\s+\w+', text))
without_oxford = len(re.findall(r'\w+,\s+\w+\s+and\s+\w+', text)) - with_oxford
if with_oxford > 0 and without_oxford > 0:
return {
"consistent": False,
"with_oxford": with_oxford,
"without_oxford": without_oxford,
"detail": f"Mixed: {with_oxford} with Oxford comma, {without_oxford} without",
}
elif with_oxford > 0:
return {"consistent": True, "style": "Oxford comma used", "count": with_oxford}
elif without_oxford > 0:
return {"consistent": True, "style": "No Oxford comma", "count": without_oxford}
return {"consistent": True, "style": "No instances to check", "count": 0}
def check_heading_case(text):
"""Check heading capitalization consistency."""
headings = re.findall(r'^#{1,6}\s+(.+)', text, re.MULTILINE)
if len(headings) < 2:
return {"consistent": True, "detail": "Not enough headings to check"}
title_case = 0
sentence_case = 0
for h in headings:
words = h.split()
if len(words) < 2:
continue
# Check if non-first words are capitalized (title case signal)
caps = sum(1 for w in words[1:] if w[0].isupper() and w.lower() not in
['a', 'an', 'the', 'and', 'or', 'but', 'in', 'on', 'at', 'to', 'for', 'of', 'with', 'by'])
if caps > len(words[1:]) * 0.5:
title_case += 1
else:
sentence_case += 1
if title_case > 0 and sentence_case > 0:
return {
"consistent": False,
"title_case": title_case,
"sentence_case": sentence_case,
"detail": f"Mixed: {title_case} Title Case, {sentence_case} sentence case",
}
return {
"consistent": True,
"style": "Title Case" if title_case > sentence_case else "Sentence case",
}
def check_contractions(text):
"""Check contraction consistency."""
contractions = len(re.findall(r"\b\w+'(t|s|re|ve|ll|d|m)\b", text))
full_forms = len(re.findall(
r'\b(do not|does not|is not|are not|was not|were not|will not|'
r'would not|could not|should not|cannot|have not|has not|'
r'it is|he is|she is|that is|there is|we are|they are|you are)\b',
text, re.IGNORECASE
))
if contractions > 0 and full_forms > 0:
return {
"consistent": False,
"contractions": contractions,
"full_forms": full_forms,
"detail": f"Mixed: {contractions} contractions, {full_forms} full forms",
}
return {
"consistent": True,
"style": "Contractions" if contractions > full_forms else "Full forms",
"contractions": contractions,
"full_forms": full_forms,
}
def check_number_format(text):
"""Check number formatting consistency."""
# Spelled out small numbers vs digits
spelled = len(re.findall(r'\b(one|two|three|four|five|six|seven|eight|nine)\b', text, re.I))
digits_small = len(re.findall(r'\b[1-9]\b', text)) # Single digits as numbers
issues = []
if spelled > 0 and digits_small > 0:
issues.append(f"Mixed: {spelled} spelled out, {digits_small} as digits for 1-9")
return {
"consistent": len(issues) == 0,
"spelled_small": spelled,
"digit_small": digits_small,
"issues": issues,
}
def check_list_punctuation(text):
"""Check bullet list punctuation consistency."""
list_items = re.findall(r'^\s*[\-\*]\s+(.+)$', text, re.MULTILINE)
if len(list_items) < 3:
return {"consistent": True, "detail": "Not enough list items"}
with_period = sum(1 for item in list_items if item.strip().endswith('.'))
without_period = len(list_items) - with_period
if with_period > 0 and without_period > 0:
return {
"consistent": False,
"with_period": with_period,
"without_period": without_period,
"detail": f"Mixed: {with_period} with period, {without_period} without",
}
return {
"consistent": True,
"style": "With periods" if with_period > without_period else "Without periods",
}
def check_quote_style(text):
"""Check quotation mark consistency."""
curly = len(re.findall(r'[\u201c\u201d\u2018\u2019]', text))
straight = len(re.findall(r'(?<!\w)["\'](?!\w)', text))
if curly > 0 and straight > 0:
return {
"consistent": False,
"curly": curly,
"straight": straight,
"detail": f"Mixed: {curly} curly quotes, {straight} straight quotes",
}
return {"consistent": True, "style": "Curly" if curly > straight else "Straight"}
def check_dash_usage(text):
"""Check dash consistency (em-dash, en-dash, hyphen)."""
em_proper = text.count('\u2014') # —
em_double = len(re.findall(r'(?<!\-)\-\-(?!\-)', text)) # --
en_dash = text.count('\u2013') # –
spaced_em = len(re.findall(r'\s\u2014\s', text))
unspaced_em = em_proper - spaced_em
issues = []
if em_proper > 0 and em_double > 0:
issues.append(f"Mixed em-dash styles: {em_proper} proper (—), {em_double} double-hyphen (--)")
if spaced_em > 0 and unspaced_em > 0:
issues.append(f"Mixed em-dash spacing: {spaced_em} spaced, {unspaced_em} unspaced")
return {
"consistent": len(issues) == 0,
"em_dashes": em_proper,
"double_hyphens": em_double,
"issues": issues,
}
def analyze_file(filepath):
"""Run all style checks on a file."""
text = filepath.read_text(encoding="utf-8", errors="replace")
checks = {
"oxford_comma": check_oxford_comma(text),
"heading_case": check_heading_case(text),
"contractions": check_contractions(text),
"number_format": check_number_format(text),
"list_punctuation": check_list_punctuation(text),
"quote_style": check_quote_style(text),
"dash_usage": check_dash_usage(text),
}
inconsistencies = sum(1 for c in checks.values() if not c.get("consistent", True))
score = round((1 - inconsistencies / len(checks)) * 100)
return {
"file": str(filepath),
"word_count": len(text.split()),
"score": score,
"inconsistencies": inconsistencies,
"checks": checks,
}
def main():
parser = argparse.ArgumentParser(description="Check style consistency")
group = parser.add_mutually_exclusive_group(required=True)
group.add_argument("--file", help="Single file to check")
group.add_argument("--files", nargs="+", help="Multiple files to check")
parser.add_argument("--json", action="store_true")
parser.add_argument("--verbose", action="store_true")
args = parser.parse_args()
files = [Path(args.file)] if args.file else [Path(f) for f in args.files]
results = []
for fp in files:
if not fp.exists():
print(f"Warning: {fp} not found, skipping", file=sys.stderr)
continue
results.append(analyze_file(fp))
if args.json:
print(json.dumps({"files": results, "total": len(results)}, indent=2))
else:
for r in results:
print(f"\n{'='*55}")
print(f" STYLE CHECK: {r['file']}")
print(f" Score: {r['score']}/100 | Inconsistencies: {r['inconsistencies']}")
print(f"{'='*55}")
for name, check in r["checks"].items():
status = "PASS" if check.get("consistent", True) else "FAIL"
detail = check.get("detail", check.get("style", ""))
issues = check.get("issues", [])
print(f" [{status}] {name}: {detail}")
if args.verbose and issues:
for issue in issues:
print(f" {issue}")
if len(results) > 1:
avg = round(sum(r["score"] for r in results) / len(results))
print(f"\n Average consistency: {avg}/100 across {len(results)} files")
print()
if __name__ == "__main__":
main()
Related skills
FAQ
What is the core method?
The Seven Sweeps Framework, seven sequential passes each focused on one dimension.
Does it check claims?
Yes, it includes a fact-checking protocol to validate proof and claims.