Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
aws-samples avatar

Browser Automation

  • 39 installs
  • 186 repo stars
  • Updated August 4, 2026
  • aws-samples/sample-strands-agent-with-agentcore

browser-automation is an agent skill that automates web browser UI interaction, form-filling, and scraping via natural language.

About

browser-automation lets an agent control a web browser for tasks needing UI interaction, login-protected pages, or human-like browsing when APIs are insufficient. A developer uses it to fill forms, click through multi-step flows, and scrape rendered pages. It exposes natural-language browser actions plus fast DOM inspection and screenshot tools.

  • Drives a browser with natural-language actions (click, type, scroll)
  • Extracts page structure, text, tables, and links without AI
  • Handles login-protected and JS-heavy pages APIs cannot reach

Browser Automation by the numbers

  • 39 all-time installs (skills.sh)
  • Ranked #1,147 of 2,715 Automation & Workflows skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

browser-automation capabilities & compatibility

Free; uses a bundled browser toolset, no external keys stated.

Capabilities
browser automation · web scraping · form filling · page extraction
Use cases
web scraping · web search · testing
Pricing
Free
From the docs

What browser-automation says it does

Web browser automation for tasks requiring UI interaction, login-protected pages, or human-like browsing when APIs are insufficient.
SKILL.md
npx skills add https://github.com/aws-samples/sample-strands-agent-with-agentcore --skill browser-automation

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs39
repo stars186
Last updatedAugust 4, 2026
Repositoryaws-samples/sample-strands-agent-with-agentcore

What it does

Automate browser UI interactions, form-filling, and scraping of login-protected or JS-heavy pages via natural language.

Who is it for?

Automating multi-step web UI flows and scraping pages that plain HTTP or APIs cannot reach.

Skip if: General information lookup on public pages, where web search or url_fetcher is preferred as lighter and faster.

When should I use this skill?

A task needs form-filling, login-gated content, JS-rendered pages, or human-like browsing.

What you get

Completed browser interactions and extracted page data (text, tables, links, screenshots).

By the numbers

  • 4 tools (browser_act, browser_get_page_info, browser_manage_tabs, browser_save_screenshot)
  • page info under 300ms
  • up to 200 links extracted

Files

SKILL.mdMarkdownGitHub ↗

Browser Automation

Available Tools

  • browser_act(instruction, starting_url?): Execute browser actions using natural language (click, type, scroll, select). Use starting_url to navigate to a page and act in a single call.
  • browser_get_page_info(url?, text?, tables?, links?): Get page structure and DOM data (fast, no AI). Use url to navigate first; text=True for full text, tables=True for table data, links=True for all links.
  • browser_manage_tabs(action, tab_index?, url?): Switch, close, or create browser tabs
  • browser_save_screenshot(filename): Save current page screenshot to workspace

When to Use

Use browser automation when the task genuinely requires it:

  • UI interactions: Filling forms, clicking buttons, navigating multi-step workflows
  • Login-required pages: Accessing content behind authentication that APIs cannot reach
  • Dynamic/JS-heavy pages: Content rendered client-side that plain HTTP requests can't capture
  • Human-like browsing needed: Sites that block bots or require realistic interaction patterns
  • Scraping structured data: When no API exists and the data must be extracted from rendered pages

Prefer web search or url_fetcher for general information lookup, news, or publicly accessible pages — browser automation is slower and heavier. Reserve it for tasks where simpler tools are insufficient.

Tool Selection

  • browser_act: UI interactions (click, type, scroll, form fill). Use starting_url to open a page and act in one call.
  • browser_get_page_info: Fast page structure check and optional content extraction (<300ms). Use url to navigate first.
  • browser_manage_tabs: Switch/close/create tabs (view tabs via get_page_info)
  • browser_save_screenshot: Save milestone screenshots (search results, confirmations, key data)

browser_act Best Practice

  • Combine up to 3 predictable steps: "1. Type 'laptop' in search 2. Click search button 3. Click first result"
  • Use starting_url when opening a fresh page: browser_act(instruction='Search for laptops', starting_url='https://amazon.com')
  • On failure: check the screenshot to see current state, then retry from that point
  • For visual creation (diagrams, drawings), prefer code/text input methods over mouse interactions

browser_get_page_info Best Practice

  • Use url to navigate and inspect in one call: browser_get_page_info(url='https://example.com', tables=True)
  • Use text=True to get full page text content (useful for reading article text)
  • Use tables=True to extract structured table data from the page
  • Use links=True to get all links on the page (up to 200)

UI Guidance (from tools-config)

Tool Selection:

  • browser_act: UI interactions (click, type, scroll, form fill). Use starting_url to navigate and act in one call.
  • browser_get_page_info: Fast DOM inspection (<300ms). Use url param to navigate first; text/tables/links params for content extraction.
  • browser_manage_tabs: Switch, close, or create tabs.
  • browser_save_screenshot: Save milestone screenshots to workspace for documents.

browser_act Best Practice:

  • Combine up to 3 predictable steps: "1. Type 'laptop' in search 2. Click search button 3. Click first result"
  • Use starting_url when opening a fresh page: browser_act(instruction='...', starting_url='https://...')
  • On failure: check the screenshot to see current state, then retry from that point

browser_get_page_info Best Practice:

  • Use url param to navigate and inspect in one call: browser_get_page_info(url='https://...', tables=True)
  • Use text=True for full page text, tables=True for table data, links=True for all page links

Related skills

FAQ

When should I not use browser automation?

For general info lookup on public pages, prefer web search or url_fetcher since browser automation is slower and heavier.

How fast is page inspection?

browser_get_page_info is a fast, no-AI DOM check that runs in under 300ms.

Automation & Workflowsintegrationstesting

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.