
Article Extractor
- 458 installs
- 511 repo stars
- Updated March 11, 2026
- michalparkola/tapestry-skills-for-claude-code
article-extractor is a Claude Code skill that extracts full article text and metadata from URLs during competitor, market, or content research and saves clean readable text without ads, navigation, or clutter.
About
article-extractor is a skill from michalparkola/tapestry-skills-for-claude-code that downloads and cleans main content from blog posts, articles, and tutorials during Claude Code sessions. The skill uses Bash and Write tools to fetch URLs, strip navigation, ads, newsletter signups, and sidebar noise, and save readable plain text for downstream analysis. Developers reach for article-extractor when researching competitors, summarizing technical posts, or archiving reference material without manual copy-paste from cluttered pages. The skill activates on requests to download, extract, or save article content from a specific URL. article-extractor fits research-heavy workflows inside the terminal where clean source text improves summarization, citation, and comparison work.
- Fetches readable article body from live URLs
- Returns structured metadata for agent workflows
- Eliminates manual copy-paste during web research
- Supports competitor and market intelligence gathering
- Tapestry skill tuned for Claude Code sessions
Article Extractor by the numbers
- 458 all-time installs (skills.sh)
- Ranked #1,841 of 16,556 AI & Agent Building skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/michalparkola/tapestry-skills-for-claude-code --skill article-extractorAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 458 |
|---|---|
| repo stars | ★ 511 |
| Last updated | March 11, 2026 |
| Repository | michalparkola/tapestry-skills-for-claude-code ↗ |
How do you extract clean article text from a URL?
Extract full article text and metadata from URLs during competitor, market, or content research inside Claude Code sessions.
Who is it for?
Developers doing competitor or content research in Claude Code who need clean article text saved from blog and tutorial URLs.
Skip if: Teams needing full browser automation, paywalled scraping, or PDF extraction should skip article-extractor.
When should I use this skill?
The user provides an article or blog URL and asks to download, extract, or save the main text content.
What you get
Clean plain-text article files with main content stripped of ads, navigation, and sidebar clutter.
- Clean plain-text article files
By the numbers
- Uses Bash and Write as allowed Claude Code tools
Files
Article Extractor
This skill extracts the main content from web articles and blog posts, removing navigation, ads, newsletter signups, and other clutter. Saves clean, readable text.
When to Use This Skill
Activate when the user:
- Provides an article/blog URL and wants the text content
- Asks to "download this article"
- Wants to "extract the content from [URL]"
- Asks to "save this blog post as text"
- Needs clean article text without distractions
How It Works
Priority Order:
1. Check if tools are installed (reader or trafilatura) 2. Download and extract article using best available tool 3. Clean up the content (remove extra whitespace, format properly) 4. Save to file with article title as filename 5. Confirm location and show preview
Installation Check
Check for article extraction tools in this order:
Option 1: reader (Recommended - Mozilla's Readability)
command -v readerIf not installed:
npm install -g @mozilla/readability-cli
# or
npm install -g reader-cliOption 2: trafilatura (Python-based, very good)
command -v trafilaturaIf not installed:
pip3 install trafilaturaOption 3: Fallback (curl + simple parsing)
If no tools available, use basic curl + text extraction (less reliable but works)
Extraction Methods
Method 1: Using reader (Best for most articles)
# Extract article
reader "URL" > article.txtPros:
- Based on Mozilla's Readability algorithm
- Excellent at removing clutter
- Preserves article structure
Method 2: Using trafilatura (Best for blogs/news)
# Extract article
trafilatura --URL "URL" --output-format txt > article.txt
# Or with more options
trafilatura --URL "URL" --output-format txt --no-comments --no-tables > article.txtPros:
- Very accurate extraction
- Good with various site structures
- Handles multiple languages
Options:
--no-comments: Skip comment sections--no-tables: Skip data tables--precision: Favor precision over recall--recall: Extract more content (may include some noise)
Method 3: Fallback (curl + basic parsing)
# Download and extract basic content
curl -s "URL" | python3 -c "
from html.parser import HTMLParser
import sys
class ArticleExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.in_content = False
self.content = []
self.skip_tags = {'script', 'style', 'nav', 'header', 'footer', 'aside'}
self.current_tag = None
def handle_starttag(self, tag, attrs):
if tag not in self.skip_tags:
if tag in {'p', 'article', 'main', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6'}:
self.in_content = True
self.current_tag = tag
def handle_data(self, data):
if self.in_content and data.strip():
self.content.append(data.strip())
def get_content(self):
return '\n\n'.join(self.content)
parser = ArticleExtractor()
parser.feed(sys.stdin.read())
print(parser.get_content())
" > article.txtNote: This is less reliable but works without dependencies.
Getting Article Title
Extract title for filename:
Using reader:
# reader outputs markdown with title at top
TITLE=$(reader "URL" | head -n 1 | sed 's/^# //')Using trafilatura:
# Get metadata including title
TITLE=$(trafilatura --URL "URL" --json | python3 -c "import json, sys; print(json.load(sys.stdin)['title'])")Using curl (fallback):
TITLE=$(curl -s "URL" | grep -oP '<title>\K[^<]+' | sed 's/ - .*//' | sed 's/ | .*//')Filename Creation
Clean title for filesystem:
# Get title
TITLE="Article Title from Website"
# Clean for filesystem (remove special chars, limit length)
FILENAME=$(echo "$TITLE" | tr '/' '-' | tr ':' '-' | tr '?' '' | tr '"' '' | tr '<' '' | tr '>' '' | tr '|' '-' | cut -c 1-100 | sed 's/ *$//')
# Add extension
FILENAME="${FILENAME}.txt"Complete Workflow
ARTICLE_URL="https://example.com/article"
# Check for tools
if command -v reader &> /dev/null; then
TOOL="reader"
echo "Using reader (Mozilla Readability)"
elif command -v trafilatura &> /dev/null; then
TOOL="trafilatura"
echo "Using trafilatura"
else
TOOL="fallback"
echo "Using fallback method (may be less accurate)"
fi
# Extract article
case $TOOL in
reader)
# Get content
reader "$ARTICLE_URL" > temp_article.txt
# Get title (first line after # in markdown)
TITLE=$(head -n 1 temp_article.txt | sed 's/^# //')
;;
trafilatura)
# Get title from metadata
METADATA=$(trafilatura --URL "$ARTICLE_URL" --json)
TITLE=$(echo "$METADATA" | python3 -c "import json, sys; print(json.load(sys.stdin).get('title', 'Article'))")
# Get clean content
trafilatura --URL "$ARTICLE_URL" --output-format txt --no-comments > temp_article.txt
;;
fallback)
# Get title
TITLE=$(curl -s "$ARTICLE_URL" | grep -oP '<title>\K[^<]+' | head -n 1)
TITLE=${TITLE%% - *} # Remove site name
TITLE=${TITLE%% | *} # Remove site name (alternate)
# Get content (basic extraction)
curl -s "$ARTICLE_URL" | python3 -c "
from html.parser import HTMLParser
import sys
class ArticleExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.in_content = False
self.content = []
self.skip_tags = {'script', 'style', 'nav', 'header', 'footer', 'aside', 'form'}
def handle_starttag(self, tag, attrs):
if tag not in self.skip_tags:
if tag in {'p', 'article', 'main'}:
self.in_content = True
if tag in {'h1', 'h2', 'h3'}:
self.content.append('\n')
def handle_data(self, data):
if self.in_content and data.strip():
self.content.append(data.strip())
def get_content(self):
return '\n\n'.join(self.content)
parser = ArticleExtractor()
parser.feed(sys.stdin.read())
print(parser.get_content())
" > temp_article.txt
;;
esac
# Clean filename
FILENAME=$(echo "$TITLE" | tr '/' '-' | tr ':' '-' | tr '?' '' | tr '"' '' | tr '<>' '' | tr '|' '-' | cut -c 1-80 | sed 's/ *$//' | sed 's/^ *//')
FILENAME="${FILENAME}.txt"
# Move to final filename
mv temp_article.txt "$FILENAME"
# Show result
echo "✓ Extracted article: $TITLE"
echo "✓ Saved to: $FILENAME"
echo ""
echo "Preview (first 10 lines):"
head -n 10 "$FILENAME"Error Handling
Common Issues
1. Tool not installed
- Try alternate tool (reader → trafilatura → fallback)
- Offer to install: "Install reader with: npm install -g reader-cli"
2. Paywall or login required
- Extraction tools may fail
- Inform user: "This article requires authentication. Cannot extract."
3. Invalid URL
- Check URL format
- Try with and without redirects
4. No content extracted
- Site may use heavy JavaScript
- Try fallback method
- Inform user if extraction fails
5. Special characters in title
- Clean title for filesystem
- Remove:
/,:,?,",<,>,| - Replace with
-or remove
Output Format
Saved File Contains:
- Article title (if available)
- Author (if available from tool)
- Main article text
- Section headings
- No navigation, ads, or clutter
What Gets Removed:
- Navigation menus
- Ads and promotional content
- Newsletter signup forms
- Related articles sidebars
- Comment sections (optional)
- Social media buttons
- Cookie notices
Tips for Best Results
1. Use reader for most articles
- Best all-around tool
- Based on Firefox Reader View
- Works on most news sites and blogs
2. Use trafilatura for:
- Academic articles
- News sites
- Blogs with complex layouts
- Non-English content
3. Fallback method limitations:
- May include some noise
- Less accurate paragraph detection
- Better than nothing for simple sites
4. Check extraction quality:
- Always show preview to user
- Ask if it looks correct
- Offer to try different tool if needed
Example Usage
Simple extraction:
# User: "Extract https://example.com/article"
reader "https://example.com/article" > temp.txt
TITLE=$(head -n 1 temp.txt | sed 's/^# //')
FILENAME="$(echo "$TITLE" | tr '/' '-').txt"
mv temp.txt "$FILENAME"
echo "✓ Saved to: $FILENAME"With error handling:
if ! reader "$URL" > temp.txt 2>/dev/null; then
if command -v trafilatura &> /dev/null; then
trafilatura --URL "$URL" --output-format txt > temp.txt
else
echo "Error: Could not extract article. Install reader or trafilatura."
exit 1
fi
fiBest Practices
- ✅ Always show preview after extraction (first 10 lines)
- ✅ Verify extraction succeeded before saving
- ✅ Clean filename for filesystem compatibility
- ✅ Try fallback method if primary fails
- ✅ Inform user which tool was used
- ✅ Keep filename length reasonable (< 100 chars)
After Extraction
Display to user: 1. "✓ Extracted: [Article Title]" 2. "✓ Saved to: [filename]" 3. Show preview (first 10-15 lines) 4. File size and location
Ask if needed:
- "Would you like me to also create a Ship-Learn-Next plan from this?" (if using ship-learn-next skill)
- "Should I extract another article?"
Related skills
FAQ
What content does article-extractor clean from pages?
article-extractor removes navigation, ads, newsletter signups, and sidebar clutter while keeping main blog or article body text. Developers use it in Claude Code when research needs readable source text instead of raw HTML.
Which Claude Code tools does article-extractor use?
article-extractor is configured with Bash and Write tools to fetch URLs and save extracted content as readable text files. The skill activates when users ask to download, extract, or save article content from a provided URL.