
Web Scraping
- 1 installs
- 5 repo stars
- Updated March 10, 2026
- ofershap/mcp-server-scraper
Extracts clean readable content, links, and metadata from URLs over MCP using Mozilla Readability, with batch scraping and in-page search.
About
This skill scrapes clean Markdown text, links, and metadata from URLs using the Readability engine, and can batch-scrape multiple pages. A developer uses it to read documentation, blogs, and articles, though it does not execute JavaScript-heavy SPAs.
- Readability-powered clean text extraction
- No headless browser, so SPAs are unsupported
Web Scraping by the numbers
- 1 all-time installs (skills.sh)
- Ranked #1,980 of 2,715 Automation & Workflows skills by installs in the Skillselion catalog
- Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/ofershap/mcp-server-scraper --skill web-scrapingAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1 |
|---|---|
| repo stars | ★ 5 |
| Last updated | March 10, 2026 |
| Repository | ofershap/mcp-server-scraper ↗ |
What it does
Extracts clean readable content, links, and metadata from URLs over MCP using Mozilla Readability, with batch scraping and in-page search.
Files
Web Scraping via MCP
Use this skill to extract clean, readable content from any URL. Returns markdown text, links, and metadata. Free alternative to Firecrawl.
Available Tools
| Tool | What it does |
|---|---|
scrape_url | Extract clean text content from a URL (Readability-powered) |
extract_links | Get all links with href and anchor text |
extract_metadata | Get title, description, OG tags, canonical, favicon |
search_page | Search for a query string within the page content |
scrape_multiple | Batch scrape multiple URLs, get title + excerpt per URL |
Workflow
1. scrape_url for reading a single page (docs, blog post, article) 2. extract_links to discover linked resources from a page 3. extract_metadata for SEO analysis or link preview data 4. scrape_multiple to survey multiple pages at once
Key Patterns
- Uses Mozilla Readability (Firefox Reader View engine) — works best with server-rendered content
- Does NOT handle JavaScript-heavy SPAs (React apps, dashboards) — use a browser MCP for those
scrape_multiplereturns title + excerpt per URL, not full content — use for surveyingsearch_pagesearches within the extracted content, not raw HTML
Limitations
- No headless browser — won't execute JavaScript
- Best for: documentation, blogs, articles, news, wikis
- Won't work for: login-gated content, SPAs, dynamically loaded content