Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
QuartzUnit avatar

Browsegrab MCP

  • Updated April 15, 2026
  • QuartzUnit/browsegrab

io.github.ArkNill/browsegrab is a MCP server that provides token-efficient Playwright browser automation via accessibility trees and MarkGrab.

About

io.github.ArkNill/browsegrab is a Web & Browser Automation MCP server aimed at developers who run local models or watch API spend closely and still need real browser interaction during development. It combines Playwright with an accessibility tree and MarkGrab so agents see a tighter, more actionable page representation than naive HTML scraping. That makes it valuable across Validate prototypes, Build-time agent tooling, and Ship testing when you want the model to drive clicks and assertions without standing up a separate Selenium grid. It is early-stage (0.1.1) and PyPI-distributed with stdio transport—expect to own browser dependencies and sandbox policy yourself. Choose Browsegrab when browser MCP cost and fidelity matter; pair with Diffgrab when you also need structured change tracking over time.

  • Playwright-driven browser agent over stdio MCP (PyPI browsegrab v0.1.1)
  • Accessibility-tree navigation to cut token burn vs raw DOM dumps
  • MarkGrab capture integration for structured page state
  • Positioned for local LLMs and token-efficient agent loops
  • GitHub source: QuartzUnit/browsegrab

Browsegrab MCP by the numbers

  • Data as of Aug 10, 2026 (Skillselion catalog sync)
terminal
claude mcp add browsegrab -- uvx browsegrab

Add your badge

Show developers this MCP server is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Packagebrowsegrab
TransportSTDIO
AuthNone
Last updatedApril 15, 2026
RepositoryQuartzUnit/browsegrab

What it does

Equip local or cost-sensitive agents with Playwright browsing that uses accessibility trees and MarkGrab for lower token use.

Who is it for?

Best when you're automating QA or research flows with Claude Code, Cursor, or local models under tight token budgets.

Skip if: Skip if you only need static HTTP fetch, lack Playwright install appetite, or want fully managed cloud browser farms.

What you get

After install, your agent can navigate and inspect pages through lean accessibility-tree and MarkGrab outputs suited to automation loops.

  • Interactive browser sessions driven from the agent
  • Lean page representations via accessibility tree and MarkGrab
  • Reusable automation for smoke tests and research crawls

By the numbers

  • Server version 0.1.1
  • PyPI identifier browsegrab with stdio transport
  • Documented stack: Playwright + accessibility tree + MarkGrab
README.md

browsegrab

한국어 문서 · llms.txt

Token-efficient browser agent for local LLMs — Playwright + accessibility tree + MarkGrab, MCP native.

browsegrab is a lightweight browser automation library designed for local LLMs (8B-35B parameters). It combines Playwright's accessibility tree with MarkGrab's HTML-to-markdown conversion to achieve 5-8x fewer tokens per step compared to alternatives like browser-use.

Features

  • Token-efficient: ~500-1,500 tokens/step (vs 4,000-10,000 for browser-use)
  • Local LLM first: Optimized for vLLM, Ollama, and OpenAI-compatible endpoints
  • MCP native: Built-in MCP server with 8 browser automation tools
  • MarkGrab integration: HTML → clean markdown for content extraction
  • Accessibility tree + ref system: Stable element references (e1, e2, ...) without vision models
  • Success pattern caching: Zero LLM calls on repeated workflows
  • 5-stage JSON parser: Robust action parsing for local LLM outputs
  • Minimal dependencies: Only playwright + httpx in core

Installation

pip install browsegrab
playwright install chromium

With optional features:

pip install browsegrab[mcp]      # MCP server support
pip install browsegrab[content]  # MarkGrab content extraction
pip install browsegrab[cli]      # CLI with rich output
pip install browsegrab[all]      # Everything

Quick Start

Python API

from browsegrab import BrowseSession

async with BrowseSession() as session:
    # Navigate and get accessibility tree snapshot
    await session.navigate("https://example.com")
    snap = await session.snapshot()
    print(snap.tree_text)
    # - heading "Example Domain" [level=1]
    # - link "Learn more": [ref=e1]

    # Click using ref ID
    result = await session.click("e1")
    print(result.url)  # https://www.iana.org/help/example-domains

    # Type into search box
    await session.navigate("https://en.wikipedia.org")
    snap = await session.snapshot()
    await session.type("e4", "Python programming", submit=True)

    # Extract compressed content (AX tree + markdown)
    content = await session.extract_content()

CLI

# Accessibility tree snapshot
browsegrab snapshot https://example.com

# JSON output
browsegrab snapshot https://example.com -f json

# Extract content (AX tree + markdown)
browsegrab extract https://en.wikipedia.org/wiki/Python

# Agentic browse (requires LLM endpoint)
browsegrab browse https://example.com "Find the about page"

MCP Server

browsegrab-mcp  # Start MCP server (stdio)

Claude Desktop / Cursor / VS Code config:

{
  "mcpServers": {
    "browsegrab": {
      "command": "browsegrab-mcp"
    }
  }
}

8 MCP tools: browser_navigate, browser_click, browser_type, browser_snapshot, browser_scroll, browser_extract_content, browser_go_back, browser_wait

How It Works

Agent Browse Loop

flowchart LR
    A["🌐 URL + Goal"] --> B["Navigate"]
    B --> C["AX Tree Snapshot\n~200–500 tokens"]
    C --> D{"LLM\nDecision"}
    D -->|"click / type / scroll"| E["Execute Action"]
    E --> C
    D -->|"goal reached"| F["Extract Content\n(MarkGrab)"]
    F --> G["✅ Result"]

Token Efficiency

browsegrab separates structure (accessibility tree) from content (MarkGrab markdown), sending only what the LLM needs:

flowchart TD
    A["Raw HTML"] --> B["Accessibility Tree"]
    A --> C["MarkGrab Markdown"]
    B --> D["Structure: ~200–500 tokens\nInteractive elements with ref IDs"]
    C --> E["Content: ~300–800 tokens\nClean markdown · on-demand"]
    D --> F["Combined: ~500–1,300 tokens/step\n⚡ 5–8× fewer than browser-use"]
    E --> F

Token efficiency (measured)

Page Interactive elements Tokens browser-use equivalent
example.com 1 ~60 ~500+
Wikipedia article 452 ~1,254 ~10,000+

Architecture

browsegrab/
├── config.py                 # Dataclass configs (env var loading)
├── result.py                 # Result types (ActionResult, BrowseResult, ...)
├── session.py                # BrowseSession orchestrator
├── browser/
│   ├── manager.py            # Playwright lifecycle (async context manager)
│   ├── snapshot.py           # Accessibility tree + ref system
│   ├── selectors.py          # 4-strategy selector resolver
│   └── actions.py            # navigate, click, type, scroll, go_back, wait
├── dom/
│   ├── ref_map.py            # ref ID ↔ element bidirectional mapping
│   └── compress.py           # AX tree + MarkGrab → compressed context
├── llm/
│   ├── base.py               # LLMProvider ABC
│   ├── provider.py           # vLLM, Ollama, OpenAI-compatible
│   ├── prompt.py             # System prompts (~400 tokens)
│   └── parse.py              # 5-stage JSON fallback parser
├── agent/
│   ├── history.py            # Sliding window history compression
│   ├── cache.py              # Domain-based success pattern cache
│   └── loop_guard.py         # Duplicate action detection
├── __main__.py               # CLI (click)
└── mcp_server.py             # FastMCP server (8 tools)

Configuration

All settings via environment variables (BROWSEGRAB_* prefix):

# Browser
BROWSEGRAB_BROWSER_HEADLESS=true
BROWSEGRAB_BROWSER_TIMEOUT_MS=30000

# LLM (for agentic browse)
BROWSEGRAB_LLM_PROVIDER=vllm          # vllm | ollama | openai
BROWSEGRAB_LLM_BASE_URL=http://localhost:8000/v1
BROWSEGRAB_LLM_MODEL=Qwen/Qwen3.5-32B-AWQ

# Agent
BROWSEGRAB_AGENT_MAX_STEPS=10
BROWSEGRAB_AGENT_ENABLE_CACHE=true

Part of the QuartzUnit Ecosystem

Library Role
markgrab Passive extraction (URL → markdown)
snapgrab Passive capture (URL → screenshot)
docpick Document OCR → structured JSON
browsegrab Active automation (goal → browser actions → results)

Development

git clone https://github.com/QuartzUnit/browsegrab.git
cd browsegrab
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
playwright install chromium

# Unit tests (no browser needed)
pytest tests/ -m "not e2e"

# Full suite including E2E
pytest tests/ -v

License

MIT


Part of the QuartzUnit ecosystem — composable Python libraries for data collection, extraction, search, and AI agent safety.

Recommended MCP Servers

How it compares

Token-efficient Playwright MCP agent, not a static scraper skill or hosting marketplace.

FAQ

Who is io.github.ArkNill/browsegrab for?

Developers who want coding agents to control a real browser with Playwright while keeping prompts smaller via accessibility trees and MarkGrab.

When should I use io.github.ArkNill/browsegrab?

Use it during Build agent-tooling, Validate prototyping, or Ship testing whenever you need interactive browser steps instead of one-shot HTTP reads.

How do I add io.github.ArkNill/browsegrab to my agent?

Install the PyPI package browsegrab 0.1.1, ensure Playwright browsers are available, register stdio MCP in your host, then invoke Browsegrab MCP tools from your agent config.

Web & Browser Automationtestingintegrations

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.