
Web Fetch
- 946 installs
- 52 repo stars
- Updated June 24, 2026
- 0xbigboss/claude-code
web-fetch is a CLI fetch skill that lets coding agents reliably read live web pages and convert them into clean markdown using bun, linkedom, and Turndown for research, specs, or context.
About
web-fetch is a bun-based fetch script in 0xbigboss/claude-code that retrieves a URL, parses HTML with linkedom, selects content-rich elements such as article tags, and converts the result to markdown via TurndownService. Usage is `bun fetch.ts <url>` with a ClaudeCode/1.0 user agent header. Developers reach for web-fetch when agents need live documentation, blog posts, or reference pages as markdown instead of raw HTML. The pipeline covers fetch, DOM parse, content element selection, and markdown conversion for agent-readable output.
- Fetches any URL with realistic browser User-Agent
- Uses linkedom to parse real DOM
- Smart content extraction prioritizing article, main, or .content elements
- Strips navigation, scripts, sidebars and noise before conversion
- Converts cleaned HTML to markdown using Turndown
Web Fetch by the numbers
- 946 all-time installs (skills.sh)
- +5 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #305 of 2,715 Automation & Workflows skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/0xbigboss/claude-code --skill web-fetchAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 946 |
|---|---|
| repo stars | ★ 52 |
| Security audit | 2 / 3 scanners passed |
| Last updated | June 24, 2026 |
| Repository | 0xbigboss/claude-code ↗ |
How do you fetch a web page as markdown for agents?
Let their coding agent reliably read live web pages and convert them into clean markdown for research, specs, or context.
Who is it for?
Developers whose coding agents need live web documentation or articles converted to markdown during research or spec gathering.
Skip if: Authenticated pages, heavy JavaScript SPAs without static content, or teams needing full browser automation instead of HTTP fetch.
When should I use this skill?
User asks agent to read a live URL, fetch web documentation, or convert a webpage to markdown for context.
What you get
Clean markdown converted from live web page HTML via bun, linkedom, and Turndown
- markdown page content
- extracted article text from URL
Files
Web Content Fetching
Fetch web content in this order: 1. Prefer markdown-native endpoints (content-type: text/markdown) 2. Use selector-based HTML extraction for known sites 3. Use the bundled Bun fallback script when selectors fail
Prerequisites
Verify required tools before extracting:
command -v curl >/dev/null || echo "curl is required"
command -v html2markdown >/dev/null || echo "html2markdown is required for HTML extraction"
command -v bun >/dev/null || echo "bun is required for fetch.ts fallback"Install Bun dependencies for the bundled script:
cd ~/.claude/skills/web-fetch && bun installDefault Workflow
Use this as the default flow for any URL:
URL="<url>"
CONTENT_TYPE="$(curl -sIL "$URL" | awk -F': ' 'tolower($1)=="content-type"{print tolower($2)}' | tr -d '\r' | tail -1)"
if echo "$CONTENT_TYPE" | grep -q "markdown"; then
curl -sL "$URL"
else
curl -sL "$URL" \
| html2markdown \
--include-selector "article,main,[role=main]" \
--exclude-selector "nav,header,footer,script,style"
fiKnown Site Selectors
| Site | Include Selector | Exclude Selector |
|---|---|---|
| platform.claude.com | #content-container | - |
| docs.anthropic.com | #content-container | - |
| developer.mozilla.org | article | - |
| github.com (docs) | article | nav,.sidebar |
| Generic | article,main,[role=main] | nav,header,footer,script,style |
Example:
curl -sL "<url>" \
| html2markdown \
--include-selector "#content-container" \
--exclude-selector "nav,header,footer"Finding the Right Selector
When a site isn't in the patterns list:
# Check what content containers exist
curl -s "<url>" | grep -o '<article[^>]*>\|<main[^>]*>\|id="[^"]*content[^"]*"' | head -10
# Test a selector
curl -sL "<url>" | html2markdown --include-selector "<selector>" | head -30
# Check line count
curl -sL "<url>" | html2markdown --include-selector "<selector>" | wc -lUniversal Fallback Script
When selectors produce poor output, run the bundled parser:
bun ~/.claude/skills/web-fetch/fetch.ts "<url>"If already in the skill directory:
bun fetch.ts "<url>"Options Reference
--include-selector "CSS" # Keep only matching elements
--exclude-selector "CSS" # Remove matching elements
--domain "https://..." # Convert relative links to absoluteTroubleshooting
Empty output with selectors: The page might be markdown-native. Check headers first:
curl -sIL "<url>" | grep -i '^content-type:'Wrong content selected: The site may have multiple article/main regions:
curl -s "<url>" | grep -o '<article[^>]*>'`html2markdown` not found: Install it, then retry selector-based extraction.
`bun` or script deps missing: Run cd ~/.claude/skills/web-fetch && bun install.
Missing code blocks: Check if the site uses non-standard code formatting.
Client-rendered content: If HTML only has "Loading..." placeholders, the content is JS-rendered. Neither curl nor the Bun script can extract it; use browser-based tools.
node_modules/
*.lock
import { parseHTML } from "linkedom";
import TurndownService from "turndown";
const url = process.argv[2];
if (!url) {
console.error("Usage: bun fetch.ts <url>");
process.exit(1);
}
// Step 1: Fetch
const response = await fetch(url, {
headers: {
"User-Agent": "Mozilla/5.0 (compatible; ClaudeCode/1.0)",
},
});
if (!response.ok) {
console.error(`Fetch failed: ${response.status} ${response.statusText}`);
process.exit(1);
}
const html = await response.text();
// Step 2: Parse DOM
const { document } = parseHTML(html);
// Step 3: Find the content-rich element
const candidates = [
...document.querySelectorAll("article"),
...document.querySelectorAll("main"),
...document.querySelectorAll('[role="main"]'),
...document.querySelectorAll(".content"),
...document.querySelectorAll("#content"),
];
let contentEl: Element | null = null;
let maxLength = 0;
for (const el of candidates) {
const len = el.textContent?.length || 0;
if (len > maxLength) {
maxLength = len;
contentEl = el;
}
}
if (!contentEl) {
contentEl = document.body;
}
// Step 4: Clean up the content element before conversion
// Remove navigation elements
const removeSelectors = [
"nav",
"header",
"footer",
"script",
"style",
"noscript",
'[role="navigation"]',
".sidebar",
".nav",
".menu",
".toc",
'[aria-label="breadcrumb"]',
];
for (const selector of removeSelectors) {
contentEl.querySelectorAll(selector).forEach((el) => el.remove());
}
// Step 5: Convert to Markdown with Turndown
const turndown = new TurndownService({
headingStyle: "atx",
codeBlockStyle: "fenced",
});
// Better code block handling
turndown.addRule("fencedCodeBlock", {
filter: (node) => {
return (
node.nodeName === "PRE" &&
node.firstChild &&
node.firstChild.nodeName === "CODE"
);
},
replacement: (content, node) => {
const el = node as Element;
const code = el.querySelector("code");
const className = code?.className || "";
const lang = className.match(/language-(\w+)/)?.[1] || "";
const text = code?.textContent || "";
return `\n\`\`\`${lang}\n${text}\n\`\`\`\n`;
},
});
// Handle pre without code child
turndown.addRule("preBlock", {
filter: (node) => {
return (
node.nodeName === "PRE" &&
(!node.firstChild || node.firstChild.nodeName !== "CODE")
);
},
replacement: (content, node) => {
const text = (node as Element).textContent || "";
return `\n\`\`\`\n${text}\n\`\`\`\n`;
},
});
// Remove "Copy page" buttons and similar UI elements
turndown.addRule("removeButtons", {
filter: (node) => {
if (node.nodeName === "BUTTON") return true;
const el = node as Element;
if (el.getAttribute?.("aria-label")?.includes("Copy")) return true;
return false;
},
replacement: () => "",
});
const markdown = turndown.turndown(contentEl.innerHTML);
// Step 6: Clean up the output
const cleaned = markdown
// Remove Loading... placeholders
.replace(/^Loading\.\.\.$/gm, "")
// Remove Copy buttons
.replace(/^Copy page$/gm, "")
.replace(/^Copy$/gm, "")
// Fix empty headings (## \n\nActual heading -> ## Actual heading)
.replace(/^(#{1,6})\s*\n\n+([A-Z])/gm, "$1 $2")
// Remove completely empty headings
.replace(/^#{1,6}\s*$/gm, "")
// Collapse multiple newlines
.replace(/\n{3,}/g, "\n\n")
.trim();
// Output with title
const title = document.title || "Untitled";
console.log(`# ${title}\n`);
console.log(cleaned);
{
"name": "web-fetch",
"version": "1.0.0",
"type": "module",
"description": "Intelligent web content extraction to clean markdown",
"dependencies": {
"linkedom": "^0.18.0",
"turndown": "^7.2.0"
}
}
Related skills
How it compares
Use web-fetch for quick URL-to-markdown ingestion; use browser automation MCPs when pages require JavaScript rendering or login flows.
FAQ
How do you run web-fetch?
web-fetch runs as `bun fetch.ts <url>` from the skill’s script. The tool fetches the page with a ClaudeCode/1.0 user agent, parses HTML with linkedom, selects content elements, and converts output to markdown via Turndown.
What libraries does web-fetch use?
web-fetch uses bun for HTTP fetch, linkedom for HTML DOM parsing, and TurndownService for markdown conversion. The pipeline targets content-rich elements such as article tags before converting to agent-readable markdown.
Is Web Fetch safe to install?
skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.