
Ai Seo
- 88 installs
- 451 repo stars
- Updated July 21, 2026
- borghei/claude-skills
ai-seo is a Claude skill that optimizes content for AI search engines through generative engine optimization (GEO) so pages get cited by ChatGPT, Perplexity, and AI Overviews.
About
ai-seo is a Claude skill for generative engine optimization (GEO), getting content cited by AI search platforms rather than only ranked in traditional results. A developer or marketer uses it to run AI-visibility audits, restructure pages into self-contained extractable blocks, add schema markup, and configure robots.txt for AI crawlers. It organizes the work around structure, authority, and presence.
- Generative engine optimization (GEO) to get content cited by AI Overviews, ChatGPT, Perplexity, Claude, Gemini
- Three pillars: extractable structure, citable authority, discoverable presence
- Covers AI citability audits, schema markup, and AI-bot access configuration
Ai Seo by the numbers
- 88 all-time installs (skills.sh)
- Ranked #1,164 of 1,879 Marketing & SEO skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
ai-seo capabilities & compatibility
Free; guidance-based skill with no API keys required.
- Capabilities
- seo audit · aeo optimization · schema markup
- Use cases
- seo · marketing
- Pricing
- Free
What ai-seo says it does
Generative engine optimization (GEO) for getting cited by AI search platforms — not just ranked in traditional results.
Traditional SEO gets your page ranked. AI SEO gets your content cited. These are different optimization targets.
AI systems pull content in chunks. They find the paragraph, list, or definition that directly answers a query and extract it.
npx skills add https://github.com/borghei/claude-skills --skill ai-seoAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 88 |
|---|---|
| repo stars | ★ 451 |
| Last updated | July 21, 2026 |
| Repository | borghei/claude-skills ↗ |
What it does
Optimize content to be cited by AI search engines like AI Overviews, ChatGPT, and Perplexity via GEO techniques.
Who is it for?
Developers and marketers who want their pages cited in AI-generated answers, not just ranked in Google.
Skip if: Traditional keyword-density and backlink-only SEO where AI citation is not a goal.
When should I use this skill?
When optimizing for AI search, AI overviews, GEO, LLM visibility, Perplexity or ChatGPT citations, or entity optimization.
What you get
Pages restructured into extractable, authoritative, discoverable blocks that AI search engines cite.
- AI visibility audit
- extractable content structure
- schema markup and bot-access config
By the numbers
- three pillars of AI citability (Structure, Authority, Presence)
- audit tests top 10 target queries on Perplexity, ChatGPT, and Google AI Overviews
Files
AI SEO
Generative engine optimization (GEO) for getting cited by AI search platforms — not just ranked in traditional results.
---
Table of Contents
- Keywords
- Quick Start
- How AI Search Differs from Traditional SEO
- The Three Pillars of AI Citability
- Core Workflows
- Content Patterns That Get Cited
- Schema Markup for AI Discovery
- Bot Access Configuration
- Monitoring and Tracking
- Best Practices
- Integration Points
---
Keywords
AI SEO, generative engine optimization, GEO, AI overviews, Google SGE, ChatGPT citations, Perplexity SEO, Claude citations, AI search optimization, semantic search, entity optimization, LLM visibility, AI-generated answers, structured data, schema markup, content extractability, AI citability, GPTBot, PerplexityBot, ClaudeBot, answer engine optimization
---
Quick Start
Run an AI Visibility Audit
1. Check robots.txt for AI bot access (GPTBot, PerplexityBot, ClaudeBot) 2. Test top 10 target queries on Perplexity, ChatGPT, and Google AI Overviews 3. Document which queries cite you, which cite competitors, and what content format wins 4. Score key pages against the Extractability Checklist 5. Prioritize pages with highest gap between search volume and current AI citation presence
Optimize a Page for AI Citation
1. Add a clear definition block in the first 200 words for informational queries 2. Structure content with self-contained H2 sections that can be extracted independently 3. Add numbered steps for process queries, comparison tables for "X vs Y" queries 4. Replace all vague claims with attributed statistics ("According to [Source], [Year]") 5. Implement FAQPage, HowTo, or Article schema markup 6. Verify AI bots are allowed in robots.txt
---
How AI Search Differs from Traditional SEO
The Fundamental Shift
Traditional SEO gets your page ranked. AI SEO gets your content cited. These are different optimization targets.
| Dimension | Traditional SEO | AI SEO |
|---|---|---|
| Goal | Rank on page 1 | Get cited in AI-generated answers |
| Success metric | Click-through rate | Citation frequency |
| Content priority | Keyword density | Answer extractability |
| Authority signal | Backlinks + domain authority | Backlinks + answer quality + attribution |
| User interaction | User clicks your link | AI extracts your answer; user may never visit |
| Content format | Long-form comprehensive | Self-contained extractable blocks |
| Optimization unit | The page | The paragraph or section |
What Carries Over from Traditional SEO
- Domain authority still matters. AI systems prefer credible sources.
- Backlinks still signal trust and expertise.
- Technical SEO fundamentals (page speed, mobile-friendly, clean HTML) still apply.
- Quality content with original insights still wins.
What Changes
- Keyword density matters less than answer clarity and directness
- Page-level optimization expands to section-level and paragraph-level optimization
- Internal linking serves discoverability for AI crawlers, not just PageRank flow
- Structured data becomes a primary signal, not a nice-to-have
---
The Three Pillars of AI Citability
Pillar 1: Structure (Extractable)
AI systems pull content in chunks. They find the paragraph, list, or definition that directly answers a query and extract it. Your content must be structured so answers are self-contained.
Extractability requirements:
- Definition blocks for "what is X" queries — tight, 1-2 sentence definitions in the first 200 words
- Numbered steps for "how to do X" queries — verb-first, self-contained steps
- Comparison tables for "X vs Y" queries — clean table format with headers
- FAQ blocks for question-based queries — explicit Q&A pairs
- Statistics with full attribution for data-oriented queries
Anti-patterns that kill extractability:
- Burying the answer in paragraph 8 of a 4,000-word essay
- Requiring context from previous sections to understand any individual section
- Using narrative prose for comparisons that should be tables
- Placing key definitions only in the conclusion
Pillar 2: Authority (Citable)
AI systems do not just extract the most relevant answer — they extract the most credible one.
Authority signals in the AI era:
- Domain authority — High-DA domains get preferential citation
- Author attribution — Named authors with credentials outperform anonymous pages
- Citation chains — Your content cites credible sources, making you credible in turn
- Recency — AI systems prefer current information for time-sensitive queries
- Original data — Proprietary research, surveys, and studies get cited more because AI cannot find this data elsewhere
- Consistent entity presence — Your brand appears across authoritative sources as an entity
Pillar 3: Presence (Discoverable)
AI systems must be able to find and index your content.
Technical requirements:
- AI crawlers allowed in robots.txt
- Fast page load and clean HTML
- No JavaScript-only rendering for important content
- Schema markup for content type classification
- Proper canonical signals
- HTTPS with valid certificates
---
Core Workflows
Workflow 1: AI Visibility Audit
Step 1: Bot Access Verification
Check robots.txt for AI crawler permissions:
# These bots must NOT be blocked for AI visibility:
GPTBot # OpenAI / ChatGPT
PerplexityBot # Perplexity
ClaudeBot # Anthropic / Claude
Google-Extended # Google AI Overviews
anthropic-ai # Anthropic (alternate)
Applebot-Extended # Apple Intelligence
cohere-ai # CohereIf any AI bot is blocked, that is the single highest priority fix. Zero visibility on that platform until resolved.
Step 2: Citation Testing
Test top 10 target queries on each platform:
| Platform | How to Test | What to Record |
|---|---|---|
| Perplexity | Search at perplexity.ai, check Sources panel | Cited? Which competitors cited? Content format winning? |
| ChatGPT | Web browsing enabled, check citations | Same |
| Google AI Overviews | Google query, check AI Overview panel | Same |
| Microsoft Copilot | Search at copilot.microsoft.com, check source cards | Same |
| Claude | Web search enabled queries | Same |
Step 3: Content Extractability Scoring
Score each key page (0-7):
- [ ] Clear definition of core concept in first 200 words
- [ ] Numbered lists or step-by-step sections for process queries
- [ ] FAQ section with direct Q&A pairs
- [ ] Statistics cited with source name and year
- [ ] Comparisons in table format (not narrative)
- [ ] H1 phrased as an answer or direct statement
- [ ] Schema markup present (FAQPage, HowTo, Article)
Interpretation: 0-3 = needs major restructuring. 4-5 = good baseline. 6-7 = strong.
Step 4: Competitive Citation Analysis
For each target query, document:
- Who is currently being cited (top 3 sources per platform)
- What content format wins (definition, list, table, quote)
- What your content lacks that cited competitors provide
- Where you have unique data or expertise competitors lack
Workflow 2: Page Optimization for AI Citation
Step 1: Lead with the Answer
The first paragraph must contain the core answer to the target query. No preamble, no context-setting, no "In today's landscape..." openers.
Step 2: Structure Self-Contained Sections
Every H2 section must be answerable as a standalone excerpt:
- Each section opens with its main point
- Each section contains its own evidence
- No section requires reading previous sections to be understood
- Each section could be quoted out of context and still make sense
Step 3: Add Extractable Content Blocks
Insert 2-3 of these per key page:
- Definition block (first 200 words)
- Numbered how-to steps (5-10 max, verb-first)
- Comparison table (clean headers, structured data)
- FAQ pairs (question matches natural language query)
- Attributed statistics ("According to [Source] ([Year]), X% of...")
- Expert quote block ("[Name], [Role at Organization]: '[quote]'")
Step 4: Replace Vague with Specific
Find and replace every vague claim:
- "Many companies" → name the companies or cite the count
- "Studies show" → name the study, organization, and year
- "Significantly improved" → state the percentage improvement
- "Leading brands" → name at least one
- "Experts say" → name the expert with credentials
Step 5: Add Schema Markup
Implement JSON-LD in the page head:
| Content Type | Schema | Impact |
|---|---|---|
| FAQ sections | FAQPage | High — AI extracts Q&A pairs directly |
| Step-by-step guides | HowTo | High — AI uses step structure |
| Articles and posts | Article | Medium — establishes content authority |
| Product pages | Product | Medium — product comparison queries |
| Author pages | Person | Medium — author credibility signal |
| Company pages | Organization | Medium — entity authority |
Workflow 3: Entity Optimization
Step 1: Define Your Entity
Ensure your brand exists as a recognized entity across the web:
- Wikipedia or Wikidata presence
- Google Knowledge Panel
- Consistent NAP (name, address, phone) across citations
- Structured About page with Organization schema
Step 2: Build Entity Associations
Connect your entity to relevant topics:
- Publish original research on topics you want to be cited for
- Get mentioned (with links) on authoritative sites in your domain
- Contribute expert quotes to industry publications
- Maintain active presence on platforms AI systems index
Step 3: Strengthen the Citation Chain
Create a network of credible references:
- Your content cites authoritative sources
- Authoritative sources cite your content
- Your author pages link to credentials and publications
- Your brand appears in industry roundups and comparisons
---
Content Patterns That Get Cited
Pattern 1: Definition Block
**[Term]** is [concise definition in 1-2 sentences]. [One sentence of context
explaining why it matters or how it differs from related concepts].Place within the first 200 words. No hedging, no preamble.
Pattern 2: Numbered Steps
Requirements for AI extraction:
- Steps are numbered (not bulleted)
- Each step starts with an action verb
- Each step is self-contained (could be quoted alone)
- 5-10 steps maximum (AI truncates longer lists)
- Each step has a brief explanation (1-2 sentences)
Pattern 3: Comparison Table
Two-column or multi-column tables with clean headers:
| Dimension | Option A | Option B |
|-----------|----------|----------|
| Price | $X/mo | $Y/mo |
| Key Feature | Description | Description |
| Best For | Use case | Use case |Pattern 4: FAQ Block
Explicit Q&A pairs. Questions should match natural language queries:
### What is [topic]?
[Direct answer in 1-2 sentences.]
### How does [topic] work?
[Step-by-step explanation.]Mark up with FAQPage schema for maximum discoverability.
Pattern 5: Attributed Statistics
According to [Source Name] ([Year]), X% of [population] [finding].Complete attribution is critical. Unattributed statistics get deprioritized because AI cannot verify the source.
Pattern 6: Expert Quote Block
"[Quote]" — [Name], [Role] at [Organization]Named experts with credentials produce citable units AI systems pick up.
---
Schema Markup for AI Discovery
Priority Implementations
FAQPage Schema (highest impact for informational queries):
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "What is [topic]?",
"acceptedAnswer": {
"@type": "Answer",
"text": "[Direct answer]"
}
}
]
}HowTo Schema (high impact for process queries):
{
"@context": "https://schema.org",
"@type": "HowTo",
"name": "How to [do thing]",
"step": [
{
"@type": "HowToStep",
"name": "Step name",
"text": "Step description"
}
]
}Article Schema (medium impact, establishes authority):
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Title",
"author": {
"@type": "Person",
"name": "Author Name",
"url": "https://author-page"
},
"datePublished": "2026-01-15",
"dateModified": "2026-03-01"
}Validate all schema at schema.org/validator before deployment.
---
Bot Access Configuration
Recommended robots.txt Configuration
# Allow all AI search crawlers
User-agent: GPTBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /
User-agent: cohere-ai
Allow: /Training vs. Citation Access
Some organizations want to allow AI citation but block training. This distinction is difficult to enforce because:
- Most AI crawlers use the same bot for both indexing and training
- Blocking the bot blocks both citation and training
- There is no industry-standard mechanism to allow one and block the other
Recommendation: Allow AI bots if you want AI citation visibility. The citation benefits outweigh the training concerns for most commercial content.
---
Monitoring and Tracking
Weekly Citation Tracking (20 minutes/week)
Test top 10 target queries on Perplexity and ChatGPT:
- Were you cited? (yes/no)
- Citation rank (1st source, 2nd, 3rd)
- What text was used from your content?
- Any new competitors appearing?
Google Search Console for AI Overviews
Use the "Search type: AI Overviews" filter in Google Search Console:
- Which queries trigger AI Overview impressions for your site
- Click-through rate from AI Overviews (typically 50-70% lower than organic)
- Which pages get cited most frequently
Monthly Monitoring Checklist
| Signal | What to Check | Tool |
|---|---|---|
| Perplexity citations | Top 10 queries | Manual testing |
| ChatGPT citations | Top 10 queries | Manual testing |
| Google AI Overviews | Impressions and clicks | Google Search Console |
| Copilot citations | Top 5 queries | Manual testing |
| AI bot crawl activity | Crawl frequency and pages | Server logs / Cloudflare |
| Competitor citations | Who is getting cited for your queries | Manual testing |
| Content freshness | Date signals on key pages | Content audit |
When Citations Drop
Diagnostic checklist when you lose a citation: 1. Did robots.txt change? (Check for accidental AI bot blocks) 2. Did a competitor publish more extractable content? 3. Did your page structure change? (Restructuring can break citation patterns) 4. Did your domain authority drop? (Check backlink profile) 5. Did the query intent shift? (AI systems may reinterpret the query)
---
Best Practices
1. Optimize at the section level, not just the page level — AI extracts paragraphs and sections, not entire pages. Every H2 block should be independently citable.
2. Lead with the answer, always — The first 200 words determine whether AI systems find your content useful. Put the answer there.
3. Attribute everything — Unattributed statistics, unnamed experts, and sourceless claims reduce your citability. Name names.
4. Update quarterly — AI systems prefer recent content. Update publish dates and refresh data points every 90 days.
5. Build entity presence — The stronger your brand's entity recognition across the web, the more AI systems trust and cite you.
6. Do not choose between traditional SEO and AI SEO — They are complementary. Many optimization signals overlap. Run both.
7. Test on multiple platforms — A page cited on Perplexity may not be cited on ChatGPT. Optimize for the platforms your audience uses.
8. Monitor competitors monthly — Track who gets cited for your target queries and study what content patterns they use.
9. Avoid JavaScript-rendered content for key answers — AI crawlers may not execute JavaScript. Ensure important content is in the initial HTML.
10. Implement schema early — FAQPage and HowTo schema are quick wins with outsized impact on AI discoverability.
---
Integration Points
- SEO Specialist — Use for traditional search ranking optimization. Run AI SEO and traditional SEO in parallel.
- Content Production — Use to create the underlying content before optimizing for AI citation.
- Content Humanizer — Use after writing. AI-sounding content performs worse in AI citations — AI systems prefer credible, human-sounding writing.
- Content Strategy — Use when deciding which topics and queries to target for AI visibility.
- Marketing Analytics — Use campaign analytics tools to track the business impact of AI citation traffic.
---
Troubleshooting
| Problem | Likely Cause | Fix |
|---|---|---|
| Content not cited despite high DA | Poor extractability — answers buried in prose | Restructure with definition blocks, numbered steps, and FAQ pairs in first 200 words |
| Cited on Perplexity but not ChatGPT | Different crawling and indexing pipelines per platform | Verify bot access for all AI crawlers; test rendering without JavaScript |
| AI Overview shows competitor instead | Competitor has more extractable, better-attributed content | Audit competitor's cited content format and match or exceed specificity |
| Citation dropped after site update | Page restructure broke the extraction pattern AI was using | Compare old vs new page structure; restore extractable blocks |
| GPTBot blocked in robots.txt unknowingly | CMS update or security plugin overwrote robots.txt | Audit robots.txt after every CMS or plugin update; set up monitoring |
| Schema markup present but no rich results | Missing required fields or content-markup mismatch | Validate with Google Rich Results Test; ensure schema matches visible page content |
| AI cites your data but not your brand | Missing entity signals — no Organization schema or sameAs links | Implement Organization schema with sameAs to Wikidata, LinkedIn, and social profiles |
---
Success Criteria
- AI citation rate: Achieve citation in 30%+ of target queries across Perplexity, ChatGPT, and Google AI Overviews within 90 days of optimization
- Extractability score: Score 6-7 out of 7 on the Content Extractability Scoring checklist for all key pages
- Bot access: Zero AI crawlers blocked in robots.txt — verified monthly with automated monitoring
- Entity recognition: Brand appears in Google Knowledge Panel and is recognized as an entity on Wikidata
- Schema coverage: 100% of content pages have appropriate JSON-LD schema (Article, FAQPage, or HowTo) validated without errors
- Freshness cadence: All key pages updated within the last 90 days with current dateModified signals
- CTR from AI Overviews: Maintain organic CTR above 0.8% for queries where AI Overviews appear (benchmark: average drops to 0.61% with AI Overviews per 2026 data)
---
Scope & Limitations
In scope:
- Optimizing content structure for AI extraction and citation
- Bot access configuration and monitoring
- Schema markup implementation for AI discoverability
- Entity optimization and Knowledge Graph presence
- Citation tracking across AI search platforms
- Content pattern design (definitions, steps, tables, FAQs)
Out of scope:
- Traditional organic ranking optimization (use SEO Specialist)
- Content creation from scratch (use Content Production)
- Paid search or paid AI placement strategies
- AI model training data licensing or opt-out negotiations
- Platform-specific API integrations for automated tracking
- Social media optimization for AI-adjacent platforms
Known limitations:
- AI citation tracking is largely manual — no standardized API exists across platforms
- Citation algorithms are opaque and change frequently without notice
- Blocking AI training while allowing citation is not technically enforceable with current bot protocols
- AI Overviews reduce traditional organic CTR by approximately 42-47% (2026 benchmarks), and this cannot be fully mitigated
---
Scripts
# Analyze content for AI citability signals
python scripts/content_scorer.py page.html --json
# Simulate how content might appear in AI search results
python scripts/serp_simulator.py --query "what is cloud cost optimization" --content page.md
# Analyze keyword opportunities for AI search visibility
python scripts/keyword_analyzer.py --keywords keywords.csv --json#!/usr/bin/env python3
"""
AI Content Citability Scorer
Analyzes content for AI search citability signals including extractability,
attribution quality, schema presence, and structural patterns that AI systems
prefer when selecting sources for generated answers.
Usage:
python content_scorer.py page.html
python content_scorer.py page.md --json
python content_scorer.py page.html --verbose
"""
import argparse
import json
import re
import sys
from pathlib import Path
def count_pattern(text, pattern):
"""Count regex pattern matches in text."""
return len(re.findall(pattern, text, re.IGNORECASE))
def analyze_extractability(text, lines):
"""Score content extractability for AI systems."""
scores = {}
# Check for definition block in first 200 words
words = text.split()
first_200 = " ".join(words[:200]) if len(words) >= 200 else text
has_definition = bool(re.search(
r'\b(is|refers to|means|defined as|describes)\b.*\.',
first_200, re.IGNORECASE
))
scores["definition_in_first_200_words"] = {
"pass": has_definition,
"detail": "Clear definition found in opening" if has_definition
else "No definition block detected in first 200 words"
}
# Check for numbered lists / steps
numbered_steps = count_pattern(text, r'^\s*\d+[\.\)]\s+', )
numbered_re = re.findall(r'^\s*\d+[\.\)]\s+', text, re.MULTILINE)
num_steps = len(numbered_re)
scores["numbered_steps"] = {
"pass": num_steps >= 3,
"count": num_steps,
"detail": f"{num_steps} numbered steps found"
}
# Check for FAQ patterns (Q&A pairs)
faq_patterns = (
count_pattern(text, r'#{1,3}\s+(what|how|why|when|where|who|is|can|do|does)\b') +
count_pattern(text, r'\*\*(what|how|why|when|where|who|is|can|do|does)\b[^*]+\*\*')
)
scores["faq_patterns"] = {
"pass": faq_patterns >= 2,
"count": faq_patterns,
"detail": f"{faq_patterns} FAQ-style question headings found"
}
# Check for comparison tables
table_rows = count_pattern(text, r'^\s*\|.*\|.*\|')
table_re = re.findall(r'^\s*\|.*\|.*\|', text, re.MULTILINE)
table_count = len(table_re)
scores["comparison_tables"] = {
"pass": table_count >= 4,
"count": table_count,
"detail": f"{table_count} table rows found"
}
# Check for attributed statistics
stat_patterns = count_pattern(
text,
r'(according to|per |based on|reported by|published by|survey|study|research)\s+[A-Z].*\d'
)
scores["attributed_statistics"] = {
"pass": stat_patterns >= 2,
"count": stat_patterns,
"detail": f"{stat_patterns} attributed statistics found"
}
# Check for self-contained H2 sections
h2_sections = re.split(r'^##\s+', text, flags=re.MULTILINE)
h2_count = max(0, len(h2_sections) - 1)
scores["h2_sections"] = {
"pass": h2_count >= 3,
"count": h2_count,
"detail": f"{h2_count} H2 sections found"
}
# Check for schema markup presence
has_schema = bool(re.search(
r'(application/ld\+json|schema\.org|@type|@context)',
text, re.IGNORECASE
))
scores["schema_markup"] = {
"pass": has_schema,
"detail": "Schema markup detected" if has_schema
else "No schema markup found"
}
return scores
def analyze_authority(text):
"""Score authority signals for AI citation preference."""
scores = {}
# Author attribution
has_author = bool(re.search(
r'(author|written by|by\s+[A-Z][a-z]+\s+[A-Z][a-z]+)',
text, re.IGNORECASE
))
scores["author_attribution"] = {
"pass": has_author,
"detail": "Author attribution detected" if has_author
else "No author attribution found"
}
# Source citations
citations = count_pattern(
text,
r'(according to|source:|cited from|published in|reported by|\(\d{4}\))'
)
scores["source_citations"] = {
"pass": citations >= 3,
"count": citations,
"detail": f"{citations} source citations found"
}
# Date signals
has_dates = bool(re.search(
r'(202[4-6]|last updated|published|modified)',
text, re.IGNORECASE
))
scores["recency_signals"] = {
"pass": has_dates,
"detail": "Date/recency signals found" if has_dates
else "No recency signals detected"
}
# Vague claims (negative signal)
vague_claims = count_pattern(
text,
r'\b(many companies|studies show|experts say|leading brands|significantly improved|a growing number)\b'
)
scores["vague_claims"] = {
"pass": vague_claims == 0,
"count": vague_claims,
"detail": f"{vague_claims} vague/unattributed claims found" if vague_claims
else "No vague claims detected"
}
return scores
def analyze_ai_patterns(text):
"""Detect AI-generated content patterns that reduce citability."""
scores = {}
# AI filler words
ai_fillers = count_pattern(
text,
r'\b(delve|landscape|crucial|leverage|robust|comprehensive|holistic|'
r'foster|facilitate|utilize|furthermore|moreover|navigate|streamline|'
r'cutting-edge|game-changer|paradigm|synergy|ecosystem|empower|'
r'transformative|seamless)\b'
)
word_count = len(text.split())
ai_density = (ai_fillers / max(word_count, 1)) * 1000
scores["ai_filler_words"] = {
"pass": ai_fillers <= 3,
"count": ai_fillers,
"per_1000_words": round(ai_density, 1),
"detail": f"{ai_fillers} AI filler words detected ({round(ai_density, 1)} per 1000 words)"
}
# Em-dash overuse
em_dashes = text.count("—") + text.count("--")
em_per_500 = (em_dashes / max(word_count, 1)) * 500
scores["em_dash_overuse"] = {
"pass": em_per_500 <= 2,
"count": em_dashes,
"per_500_words": round(em_per_500, 1),
"detail": f"{em_dashes} em-dashes ({round(em_per_500, 1)} per 500 words)"
}
# Hedging chains
hedging = count_pattern(
text,
r"(it'?s important to note|it'?s worth mentioning|one might argue|"
r"in many cases|it goes without saying|needless to say)"
)
scores["hedging_chains"] = {
"pass": hedging == 0,
"count": hedging,
"detail": f"{hedging} hedging phrases found"
}
return scores
def calculate_overall_score(extractability, authority, ai_patterns):
"""Calculate weighted overall citability score 0-100."""
total = 0
max_score = 0
# Extractability: 50% weight
for key, val in extractability.items():
max_score += 50 / len(extractability)
if val["pass"]:
total += 50 / len(extractability)
# Authority: 30% weight
for key, val in authority.items():
max_score += 30 / len(authority)
if val["pass"]:
total += 30 / len(authority)
# AI patterns: 20% weight (inverted — passing means fewer AI tells)
for key, val in ai_patterns.items():
max_score += 20 / len(ai_patterns)
if val["pass"]:
total += 20 / len(ai_patterns)
return round(total, 1)
def grade_score(score):
"""Convert numeric score to letter grade."""
if score >= 85:
return "A"
elif score >= 70:
return "B"
elif score >= 55:
return "C"
elif score >= 40:
return "D"
else:
return "F"
def main():
parser = argparse.ArgumentParser(
description="Score content for AI search citability"
)
parser.add_argument("file", help="Path to content file (HTML or Markdown)")
parser.add_argument("--json", action="store_true", help="Output as JSON")
parser.add_argument(
"--verbose", action="store_true",
help="Show detailed per-check results"
)
args = parser.parse_args()
filepath = Path(args.file)
if not filepath.exists():
print(f"Error: File not found: {filepath}", file=sys.stderr)
sys.exit(1)
text = filepath.read_text(encoding="utf-8", errors="replace")
lines = text.splitlines()
word_count = len(text.split())
extractability = analyze_extractability(text, lines)
authority = analyze_authority(text)
ai_patterns = analyze_ai_patterns(text)
overall_score = calculate_overall_score(extractability, authority, ai_patterns)
grade = grade_score(overall_score)
# Build recommendations
recommendations = []
for key, val in extractability.items():
if not val["pass"]:
recommendations.append(f"Extractability: Improve {key.replace('_', ' ')}")
for key, val in authority.items():
if not val["pass"]:
recommendations.append(f"Authority: Improve {key.replace('_', ' ')}")
for key, val in ai_patterns.items():
if not val["pass"]:
recommendations.append(f"AI Patterns: Fix {key.replace('_', ' ')}")
result = {
"file": str(filepath),
"word_count": word_count,
"overall_score": overall_score,
"grade": grade,
"extractability": extractability,
"authority": authority,
"ai_patterns": ai_patterns,
"recommendations": recommendations,
}
if args.json:
print(json.dumps(result, indent=2))
else:
print(f"\n{'='*60}")
print(f" AI CITABILITY SCORE: {overall_score}/100 (Grade: {grade})")
print(f"{'='*60}")
print(f" File: {filepath}")
print(f" Word count: {word_count}")
print()
passed = sum(1 for v in extractability.values() if v["pass"])
print(f" Extractability: {passed}/{len(extractability)} checks passed")
passed = sum(1 for v in authority.values() if v["pass"])
print(f" Authority: {passed}/{len(authority)} checks passed")
passed = sum(1 for v in ai_patterns.values() if v["pass"])
print(f" AI Patterns: {passed}/{len(ai_patterns)} checks passed")
if args.verbose:
print(f"\n --- Extractability Details ---")
for key, val in extractability.items():
status = "PASS" if val["pass"] else "FAIL"
print(f" [{status}] {key}: {val['detail']}")
print(f"\n --- Authority Details ---")
for key, val in authority.items():
status = "PASS" if val["pass"] else "FAIL"
print(f" [{status}] {key}: {val['detail']}")
print(f"\n --- AI Pattern Details ---")
for key, val in ai_patterns.items():
status = "PASS" if val["pass"] else "FAIL"
print(f" [{status}] {key}: {val['detail']}")
if recommendations:
print(f"\n --- Recommendations ---")
for i, rec in enumerate(recommendations, 1):
print(f" {i}. {rec}")
print()
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
AI SEO Keyword Analyzer
Analyzes keywords for AI search optimization potential. Classifies intent,
estimates AI Overview likelihood, scores competition signals, and prioritizes
keywords by AI citation opportunity.
Usage:
python keyword_analyzer.py --keywords keywords.csv --json
python keyword_analyzer.py --keywords keywords.csv --sort ai_score
python keyword_analyzer.py --keyword "what is cloud cost optimization"
"""
import argparse
import csv
import json
import re
import sys
from pathlib import Path
# Intent classification patterns
INTENT_PATTERNS = {
"informational": [
r'\b(what is|what are|how to|how do|why is|why do|guide|tutorial|'
r'explain|definition|meaning|example|examples|overview|introduction|'
r'difference between|compared to)\b'
],
"commercial": [
r'\b(best|top|review|reviews|comparison|vs|versus|alternative|'
r'alternatives|compare|which|rating|ratings|recommended)\b'
],
"transactional": [
r'\b(buy|purchase|price|pricing|cost|discount|deal|coupon|order|'
r'subscribe|sign up|free trial|download|get started)\b'
],
"navigational": [
r'\b(login|log in|sign in|dashboard|account|support|contact|'
r'official|site|website|app|docs|documentation)\b'
],
}
# AI Overview likelihood factors
AI_OVERVIEW_TRIGGERS = [
r'\b(what is|what are|how to|how do|why|when|where|who)\b',
r'\b(definition|meaning|explain|difference|vs|versus)\b',
r'\b(best practices|steps|process|guide|tips)\b',
r'\b(statistics|data|numbers|percentage|rate)\b',
]
# Content format mapping by intent
FORMAT_RECOMMENDATIONS = {
"informational": {
"primary": "Definition block + numbered steps",
"schema": "FAQPage, HowTo, or Article",
"ai_format": "Lead with 1-2 sentence definition, then structured steps",
},
"commercial": {
"primary": "Comparison table + pros/cons lists",
"schema": "Product, Review, or FAQPage",
"ai_format": "Structured comparison with clear recommendation",
},
"transactional": {
"primary": "Product details + pricing table",
"schema": "Product with Offer",
"ai_format": "Direct answer with pricing and CTA",
},
"navigational": {
"primary": "Direct answer with link",
"schema": "Organization or WebSite",
"ai_format": "Brand-focused answer with official link",
},
}
def classify_intent(keyword):
"""Classify search intent of a keyword."""
keyword_lower = keyword.lower()
scores = {}
for intent, patterns in INTENT_PATTERNS.items():
score = 0
for pattern in patterns:
matches = len(re.findall(pattern, keyword_lower))
score += matches
scores[intent] = score
# Default to informational if no strong signal
if max(scores.values()) == 0:
return "informational"
return max(scores, key=scores.get)
def estimate_ai_overview_likelihood(keyword):
"""Estimate probability of AI Overview appearing for this query."""
keyword_lower = keyword.lower()
triggers = 0
for pattern in AI_OVERVIEW_TRIGGERS:
if re.search(pattern, keyword_lower):
triggers += 1
# Question format bonus
if keyword_lower.strip().endswith('?') or keyword_lower.startswith(('what', 'how', 'why', 'when', 'where', 'who')):
triggers += 1
# Long-tail bonus (more words = more likely informational = more AI Overview)
word_count = len(keyword.split())
if word_count >= 4:
triggers += 1
if word_count >= 6:
triggers += 1
# Normalize to 0-100
likelihood = min(triggers * 20, 100)
if likelihood >= 80:
label = "Very High"
elif likelihood >= 60:
label = "High"
elif likelihood >= 40:
label = "Medium"
elif likelihood >= 20:
label = "Low"
else:
label = "Very Low"
return {"score": likelihood, "label": label, "triggers": triggers}
def score_ai_citation_opportunity(keyword, intent, ai_likelihood):
"""Score the AI citation opportunity for a keyword."""
score = 0
# Intent scoring (informational and commercial get cited most)
intent_scores = {
"informational": 30,
"commercial": 25,
"transactional": 10,
"navigational": 5,
}
score += intent_scores.get(intent, 10)
# AI Overview likelihood contribution
score += ai_likelihood["score"] * 0.4
# Word count factor (longer queries = more specific = better citation chance)
word_count = len(keyword.split())
if 3 <= word_count <= 6:
score += 15
elif word_count > 6:
score += 10
elif word_count <= 2:
score += 5
# Question format bonus
if any(keyword.lower().startswith(w) for w in ['what', 'how', 'why', 'when', 'where', 'who']):
score += 10
return min(round(score, 1), 100)
def analyze_keyword(keyword, volume=None, difficulty=None):
"""Full analysis of a single keyword."""
intent = classify_intent(keyword)
ai_likelihood = estimate_ai_overview_likelihood(keyword)
ai_score = score_ai_citation_opportunity(keyword, intent, ai_likelihood)
format_rec = FORMAT_RECOMMENDATIONS.get(intent, FORMAT_RECOMMENDATIONS["informational"])
result = {
"keyword": keyword,
"word_count": len(keyword.split()),
"intent": intent,
"ai_overview_likelihood": ai_likelihood,
"ai_citation_score": ai_score,
"recommended_format": format_rec,
}
if volume is not None:
result["volume"] = volume
if difficulty is not None:
result["difficulty"] = difficulty
# Priority classification
if ai_score >= 70:
result["priority"] = "High"
elif ai_score >= 45:
result["priority"] = "Medium"
else:
result["priority"] = "Low"
return result
def load_keywords_from_csv(filepath):
"""Load keywords from CSV file."""
keywords = []
with open(filepath, 'r', encoding='utf-8') as f:
reader = csv.DictReader(f)
if reader.fieldnames is None:
# Try as simple list
f.seek(0)
for line in f:
kw = line.strip()
if kw:
keywords.append({"keyword": kw})
return keywords
for row in reader:
entry = {}
# Try common column names
for key in ['keyword', 'Keyword', 'query', 'Query', 'term', 'Term']:
if key in row:
entry["keyword"] = row[key]
break
if "keyword" not in entry:
# Use first column
first_key = list(row.keys())[0]
entry["keyword"] = row[first_key]
for key in ['volume', 'Volume', 'search_volume', 'Search Volume']:
if key in row and row[key]:
try:
entry["volume"] = int(row[key].replace(',', ''))
except (ValueError, AttributeError):
pass
for key in ['difficulty', 'Difficulty', 'KD', 'kd']:
if key in row and row[key]:
try:
entry["difficulty"] = int(row[key])
except (ValueError, AttributeError):
pass
if entry.get("keyword"):
keywords.append(entry)
return keywords
def main():
parser = argparse.ArgumentParser(
description="Analyze keywords for AI search optimization potential"
)
group = parser.add_mutually_exclusive_group(required=True)
group.add_argument(
"--keywords", help="Path to CSV file with keywords"
)
group.add_argument(
"--keyword", help="Single keyword to analyze"
)
parser.add_argument("--json", action="store_true", help="Output as JSON")
parser.add_argument(
"--sort", choices=["ai_score", "volume", "keyword"],
default="ai_score",
help="Sort results by (default: ai_score)"
)
args = parser.parse_args()
if args.keyword:
result = analyze_keyword(args.keyword)
if args.json:
print(json.dumps(result, indent=2))
else:
print(f"\n{'='*60}")
print(f" KEYWORD ANALYSIS: {args.keyword}")
print(f"{'='*60}")
print(f" Intent: {result['intent']}")
print(f" AI Overview Likelihood: {result['ai_overview_likelihood']['label']} ({result['ai_overview_likelihood']['score']}%)")
print(f" AI Citation Score: {result['ai_citation_score']}/100")
print(f" Priority: {result['priority']}")
print(f"\n Recommended Format:")
print(f" Content: {result['recommended_format']['primary']}")
print(f" Schema: {result['recommended_format']['schema']}")
print(f" AI Optimization: {result['recommended_format']['ai_format']}")
print()
return
# CSV mode
filepath = Path(args.keywords)
if not filepath.exists():
print(f"Error: File not found: {filepath}", file=sys.stderr)
sys.exit(1)
keyword_data = load_keywords_from_csv(filepath)
if not keyword_data:
print("Error: No keywords found in file", file=sys.stderr)
sys.exit(1)
results = []
for entry in keyword_data:
result = analyze_keyword(
entry["keyword"],
volume=entry.get("volume"),
difficulty=entry.get("difficulty"),
)
results.append(result)
# Sort
if args.sort == "ai_score":
results.sort(key=lambda x: x["ai_citation_score"], reverse=True)
elif args.sort == "volume":
results.sort(key=lambda x: x.get("volume", 0) or 0, reverse=True)
elif args.sort == "keyword":
results.sort(key=lambda x: x["keyword"])
if args.json:
print(json.dumps({"keywords": results, "total": len(results)}, indent=2))
else:
print(f"\n{'='*70}")
print(f" AI KEYWORD ANALYSIS — {len(results)} keywords")
print(f"{'='*70}")
print(f" {'Keyword':<40} {'Intent':<14} {'AI Score':<10} {'Priority'}")
print(f" {'-'*40} {'-'*14} {'-'*10} {'-'*8}")
for r in results:
print(f" {r['keyword'][:39]:<40} {r['intent']:<14} {r['ai_citation_score']:<10} {r['priority']}")
# Summary
high = sum(1 for r in results if r["priority"] == "High")
medium = sum(1 for r in results if r["priority"] == "Medium")
low = sum(1 for r in results if r["priority"] == "Low")
print(f"\n Summary: {high} High / {medium} Medium / {low} Low priority")
print()
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
AI SERP Simulator
Simulates how content might appear when extracted by AI search systems.
Identifies the most likely extractable blocks, previews citation snippets,
and scores extraction readiness per content section.
Usage:
python serp_simulator.py --content page.md --query "what is cloud cost optimization"
python serp_simulator.py --content page.md --query "how to reduce AWS costs" --json
python serp_simulator.py --content page.md --query "best cloud optimization tools" --top 5
"""
import argparse
import json
import re
import sys
from pathlib import Path
def tokenize(text):
"""Simple word tokenization."""
return re.findall(r'\b[a-z0-9]+\b', text.lower())
def extract_sections(text):
"""Split content into sections by headings."""
sections = []
current_heading = "Introduction"
current_level = 0
current_content = []
for line in text.splitlines():
heading_match = re.match(r'^(#{1,6})\s+(.+)', line)
if heading_match:
if current_content:
content_text = "\n".join(current_content).strip()
if content_text:
sections.append({
"heading": current_heading,
"level": current_level,
"content": content_text,
"word_count": len(content_text.split()),
})
current_level = len(heading_match.group(1))
current_heading = heading_match.group(2).strip()
current_content = []
else:
current_content.append(line)
if current_content:
content_text = "\n".join(current_content).strip()
if content_text:
sections.append({
"heading": current_heading,
"level": current_level,
"content": content_text,
"word_count": len(content_text.split()),
})
return sections
def extract_paragraphs(text):
"""Split content into individual paragraphs."""
paragraphs = []
for block in re.split(r'\n\s*\n', text):
block = block.strip()
if block and not block.startswith('#') and len(block.split()) >= 10:
paragraphs.append(block)
return paragraphs
def extract_lists(text):
"""Extract numbered and bulleted lists."""
lists = []
current_list = []
in_list = False
for line in text.splitlines():
is_list_item = bool(re.match(r'^\s*(\d+[\.\)]|\-|\*)\s+', line))
if is_list_item:
current_list.append(line.strip())
in_list = True
else:
if in_list and current_list:
lists.append("\n".join(current_list))
current_list = []
in_list = False
if current_list:
lists.append("\n".join(current_list))
return lists
def extract_tables(text):
"""Extract markdown tables."""
tables = []
current_table = []
in_table = False
for line in text.splitlines():
if '|' in line and line.strip().startswith('|'):
current_table.append(line.strip())
in_table = True
else:
if in_table and len(current_table) >= 3:
tables.append("\n".join(current_table))
current_table = []
in_table = False
if in_table and len(current_table) >= 3:
tables.append("\n".join(current_table))
return tables
def score_relevance(query, text):
"""Score relevance of text block to query using term overlap."""
query_tokens = set(tokenize(query))
text_tokens = tokenize(text)
text_token_set = set(text_tokens)
if not query_tokens:
return 0.0
# Term overlap
overlap = query_tokens & text_token_set
overlap_ratio = len(overlap) / len(query_tokens)
# Term frequency boost
freq_boost = 0
for token in query_tokens:
count = text_tokens.count(token)
if count > 0:
freq_boost += min(count / len(text_tokens) * 100, 5)
# Position bonus (earlier = better)
first_500 = " ".join(text.split()[:500]).lower()
position_bonus = 0
for token in query_tokens:
if token in first_500:
position_bonus += 0.1
score = (overlap_ratio * 60) + (freq_boost * 5) + (position_bonus * 10)
return min(round(score, 1), 100)
def classify_block_type(text):
"""Classify an extractable block by type."""
if re.search(r'^\s*\d+[\.\)]\s+', text, re.MULTILINE):
return "numbered_steps"
if re.search(r'^\s*[\-\*]\s+', text, re.MULTILINE):
return "bullet_list"
if '|' in text and text.strip().startswith('|'):
return "table"
if re.search(r'\b(is|refers to|means|defined as)\b.*\.', text[:300], re.IGNORECASE):
return "definition"
if '?' in text.split('\n')[0] if text.split('\n') else False:
return "faq_answer"
return "paragraph"
def generate_snippet(text, max_chars=300):
"""Generate a citation snippet from a block."""
clean = re.sub(r'\s+', ' ', text).strip()
if len(clean) <= max_chars:
return clean
# Cut at sentence boundary
truncated = clean[:max_chars]
last_period = truncated.rfind('.')
if last_period > max_chars * 0.5:
return truncated[:last_period + 1]
return truncated + "..."
def simulate_extraction(content, query, top_n=5):
"""Simulate AI extraction process on content."""
blocks = []
# Extract different block types
paragraphs = extract_paragraphs(content)
for p in paragraphs:
blocks.append({"text": p, "source": "paragraph"})
lists = extract_lists(content)
for l in lists:
blocks.append({"text": l, "source": "list"})
tables = extract_tables(content)
for t in tables:
blocks.append({"text": t, "source": "table"})
sections = extract_sections(content)
for s in sections:
if s["word_count"] >= 20:
blocks.append({"text": s["content"], "source": f"section:{s['heading']}"})
# Score and rank
scored_blocks = []
for block in blocks:
relevance = score_relevance(query, block["text"])
block_type = classify_block_type(block["text"])
word_count = len(block["text"].split())
# Type bonus
type_bonus = {
"definition": 15,
"numbered_steps": 12,
"table": 10,
"faq_answer": 12,
"bullet_list": 5,
"paragraph": 0,
}.get(block_type, 0)
# Length penalty (too short or too long)
length_penalty = 0
if word_count < 20:
length_penalty = -20
elif word_count > 500:
length_penalty = -10
final_score = min(relevance + type_bonus + length_penalty, 100)
scored_blocks.append({
"text": block["text"],
"source": block["source"],
"block_type": block_type,
"word_count": word_count,
"relevance_score": relevance,
"type_bonus": type_bonus,
"final_score": max(final_score, 0),
"snippet": generate_snippet(block["text"]),
})
# Sort by final score
scored_blocks.sort(key=lambda x: x["final_score"], reverse=True)
# Deduplicate (remove blocks with >80% text overlap)
deduplicated = []
seen_snippets = set()
for block in scored_blocks:
snippet_key = block["snippet"][:100]
if snippet_key not in seen_snippets:
deduplicated.append(block)
seen_snippets.add(snippet_key)
return deduplicated[:top_n]
def main():
parser = argparse.ArgumentParser(
description="Simulate AI search extraction from content"
)
parser.add_argument(
"--content", required=True,
help="Path to content file (Markdown or HTML)"
)
parser.add_argument(
"--query", required=True,
help="Search query to simulate"
)
parser.add_argument(
"--top", type=int, default=5,
help="Number of top extractable blocks to show (default: 5)"
)
parser.add_argument("--json", action="store_true", help="Output as JSON")
args = parser.parse_args()
filepath = Path(args.content)
if not filepath.exists():
print(f"Error: File not found: {filepath}", file=sys.stderr)
sys.exit(1)
content = filepath.read_text(encoding="utf-8", errors="replace")
results = simulate_extraction(content, args.query, args.top)
output = {
"query": args.query,
"file": str(filepath),
"total_word_count": len(content.split()),
"top_extractable_blocks": results,
}
if args.json:
print(json.dumps(output, indent=2))
else:
print(f"\n{'='*60}")
print(f" AI SERP SIMULATION")
print(f"{'='*60}")
print(f" Query: {args.query}")
print(f" File: {filepath}")
print(f" Content: {len(content.split())} words")
print()
if not results:
print(" No extractable blocks found matching the query.")
else:
for i, block in enumerate(results, 1):
print(f" --- Block #{i} (Score: {block['final_score']}/100) ---")
print(f" Type: {block['block_type']} | Words: {block['word_count']} | Source: {block['source']}")
print(f" Snippet preview:")
# Wrap snippet for display
snippet = block["snippet"]
for line in [snippet[j:j+70] for j in range(0, len(snippet), 70)]:
print(f" {line}")
print()
print(f" Total blocks analyzed: {len(results)}")
print()
if __name__ == "__main__":
main()
Related skills
FAQ
How is AI SEO different from traditional SEO?
Traditional SEO gets a page ranked; AI SEO gets content cited in AI-generated answers, optimizing at the paragraph or section level for extractability.
What are the three pillars of AI citability?
Structure (extractable), Authority (citable), and Presence (discoverable).