Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
bighardperson avatar

Playwright Scraper Skill

  • 8 installs
  • 33 repo stars
  • Updated April 26, 2026
  • bighardperson/computer-science-skills-collection

Playwright-scraper-skill is a skill that scrapes web pages with Playwright, using stealth techniques to bypass anti-bot and Cloudflare protection.

About

Playwright-scraper-skill is a Playwright-based web scraping skill with anti-bot protection. It provides a simple script for dynamic JavaScript sites and a stealth script that hides automation markers and mimics human behavior to bypass Cloudflare. A developer uses it to fetch content from sites that block basic fetchers, choosing the method by the target's anti-bot level.

  • Playwright web scraping with anti-bot stealth and Cloudflare bypass
  • Two scripts: simple for dynamic sites, stealth for protected sites
  • Configurable via env vars: user-agent, wait time, headful, screenshots

Playwright Scraper Skill by the numbers

  • 8 all-time installs (skills.sh)
  • Ranked #1,522 of 2,719 Automation & Workflows skills by installs in the Skillselion catalog
  • Data as of Jul 30, 2026 (Skillselion catalog sync)
At a glance

playwright-scraper-skill capabilities & compatibility

Capabilities
web scraping · browser automation
Works with
playwright · chrome
Use cases
web scraping · web search
From the docs

What playwright-scraper-skill says it does

A Playwright-based web scraping OpenClaw Skill with anti-bot protection.
SKILL.md
Hide automation markers (`navigator.webdriver = false`)
SKILL.md
npx skills add https://github.com/bighardperson/computer-science-skills-collection --skill playwright-scraper-skill

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs8
repo stars33
Last updatedApril 26, 2026
Repositorybighardperson/computer-science-skills-collection

What it does

Scrape dynamic or anti-bot-protected websites using Playwright simple or stealth scripts.

Who is it for?

Fetching content from dynamic or Cloudflare-protected sites that block basic fetchers

Skip if: Repositories or source-tree code search, and simple static sites where web_fetch suffices

When should I use this skill?

A target site returns 403 or a Cloudflare challenge to basic fetching

What you get

Page content returned as JSON, optionally with a screenshot and saved HTML.

  • scraped page content JSON
  • optional screenshot and HTML

By the numbers

  • 2 bundled scraper scripts
  • 100% success rate reported on Discuss.com.hk with stealth

Files

SKILL.mdMarkdownGitHub ↗

Playwright Scraper Skill

A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Choose the best approach based on the target website's anti-bot level.

---

🎯 Use Case Matrix

Target WebsiteAnti-Bot LevelRecommended MethodScript
Regular SitesLowweb_fetch toolN/A (built-in)
Dynamic SitesMediumPlaywright Simplescripts/playwright-simple.js
Cloudflare ProtectedHighPlaywright Stealthscripts/playwright-stealth.js
YouTubeSpecialdeep-scraperInstall separately
RedditSpecialreddit-scraperInstall separately

---

📦 Installation

cd playwright-scraper-skill
npm install
npx playwright install chromium

---

🚀 Quick Start

1️⃣ Simple Sites (No Anti-Bot)

Use OpenClaw's built-in web_fetch tool:

# Invoke directly in OpenClaw
Hey, fetch me the content from https://example.com

---

2️⃣ Dynamic Sites (Requires JavaScript)

Use Playwright Simple:

node scripts/playwright-simple.js "https://example.com"

Example output:

{
  "url": "https://example.com",
  "title": "Example Domain",
  "content": "...",
  "elapsedSeconds": "3.45"
}

---

3️⃣ Anti-Bot Protected Sites (Cloudflare etc.)

Use Playwright Stealth:

node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"

Features:

  • Hide automation markers (navigator.webdriver = false)
  • Realistic User-Agent (iPhone, Android)
  • Random delays to mimic human behavior
  • Screenshot and HTML saving support

---

4️⃣ YouTube Video Transcripts

Use deep-scraper (install separately):

# Install deep-scraper skill
npx clawhub install deep-scraper

# Use it
cd skills/deep-scraper
node assets/youtube_handler.js "https://www.youtube.com/watch?v=VIDEO_ID"

---

📖 Script Descriptions

scripts/playwright-simple.js

  • Use Case: Regular dynamic websites
  • Speed: Fast (3-5 seconds)
  • Anti-Bot: None
  • Output: JSON (title, content, URL)

scripts/playwright-stealth.js

  • Use Case: Sites with Cloudflare or anti-bot protection
  • Speed: Medium (5-20 seconds)
  • Anti-Bot: Medium-High (hides automation, realistic UA)
  • Output: JSON + Screenshot + HTML file
  • Verified: 100% success on Discuss.com.hk

---

🎓 Best Practices

1. Try web_fetch First

If the site doesn't have dynamic loading, use OpenClaw's web_fetch tool—it's fastest.

2. Need JavaScript? Use Playwright Simple

If you need to wait for JavaScript rendering, use playwright-simple.js.

3. Getting Blocked? Use Stealth

If you encounter 403 or Cloudflare challenges, use playwright-stealth.js.

4. Special Sites Need Specialized Skills

  • YouTube → deep-scraper
  • Reddit → reddit-scraper
  • Twitter → bird skill

---

🔧 Customization

All scripts support environment variables:

# Set screenshot path
SCREENSHOT_PATH=/path/to/screenshot.png node scripts/playwright-stealth.js URL

# Set wait time (milliseconds)
WAIT_TIME=10000 node scripts/playwright-simple.js URL

# Enable headful mode (show browser)
HEADLESS=false node scripts/playwright-stealth.js URL

# Save HTML
SAVE_HTML=true node scripts/playwright-stealth.js URL

# Custom User-Agent
USER_AGENT="Mozilla/5.0 ..." node scripts/playwright-stealth.js URL

---

📊 Performance Comparison

MethodSpeedAnti-BotSuccess Rate (Discuss.com.hk)
web_fetch⚡ Fastest❌ None0%
Playwright Simple🚀 Fast⚠️ Low20%
Playwright Stealth⏱️ Medium✅ Medium100%
Puppeteer Stealth⏱️ Medium✅ Medium-High~80%
Crawlee (deep-scraper)🐢 Slow❌ Detected0%
Chaser (Rust)⏱️ Medium❌ Detected0%

---

🛡️ Anti-Bot Techniques Summary

Lessons learned from our testing:

✅ Effective Anti-Bot Measures

1. Hide `navigator.webdriver` — Essential 2. Realistic User-Agent — Use real devices (iPhone, Android) 3. Mimic Human Behavior — Random delays, scrolling 4. Avoid Framework Signatures — Crawlee, Selenium are easily detected 5. Use `addInitScript` (Playwright) — Inject before page load

❌ Ineffective Anti-Bot Measures

1. Only changing User-Agent — Not enough 2. Using high-level frameworks (Crawlee) — More easily detected 3. Docker isolation — Doesn't help with Cloudflare

---

🔍 Troubleshooting

Issue: 403 Forbidden

Solution: Use playwright-stealth.js

Issue: Cloudflare Challenge Page

Solution: 1. Increase wait time (10-15 seconds) 2. Try headless: false (headful mode sometimes has higher success rate) 3. Consider using proxy IPs

Issue: Blank Page

Solution: 1. Increase waitForTimeout 2. Use waitUntil: 'networkidle' or 'domcontentloaded' 3. Check if login is required

---

📝 Memory & Experience

2026-02-07 Discuss.com.hk Test Conclusions

  • Pure Playwright + Stealth succeeded (5s, 200 OK)
  • ❌ Crawlee (deep-scraper) failed (403)
  • ❌ Chaser (Rust) failed (Cloudflare)
  • ❌ Puppeteer standard failed (403)

Best Solution: Pure Playwright + anti-bot techniques (framework-independent)

---

🚧 Future Improvements

  • [ ] Add proxy IP rotation
  • [ ] Implement cookie management (maintain login state)
  • [ ] Add CAPTCHA handling (2captcha / Anti-Captcha)
  • [ ] Batch scraping (parallel URLs)
  • [ ] Integration with OpenClaw's browser tool

---

📚 References

Related skills

FAQ

When should I use the stealth script?

When you hit a 403 or Cloudflare challenge; the stealth script hides automation and uses realistic user agents.

What output does it produce?

JSON with title, content, and URL, plus optional screenshot and HTML file for the stealth script.

Automation & Workflowsintegrationsbackend

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.