
Browser Use
- 124 installs
- 1 repo stars
- Updated April 3, 2026
- shawnpana/browser-use
This is a copy of browser-use by browser-use - installs and ranking accrue to the original listing.
Enable agents to drive a real browser for UI testing, authenticated flows, scraping, and multi-step web automation.
About
Teaches agents to operate real browsers for E2E checks, form submission, dynamic SPAs, and research scrapes—bridging the gap between API-only tools and human-like web interaction for validation and autonomous workflows.
- Headed and headless browser control
- DOM interaction and selectors
- Multi-step navigation flows
- Session and cookie persistence
- UI verification for agents
Browser Use by the numbers
- 124 all-time installs (skills.sh)
- +1 installs in the week ending Jul 27, 2026 (Skillselion tracking)
- Data as of Jul 27, 2026 (Skillselion catalog sync)
npx skills add https://github.com/shawnpana/browser-use --skill browser-useAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 124 |
|---|---|
| repo stars | ★ 1 |
| Last updated | April 3, 2026 |
| Repository | shawnpana/browser-use ↗ |
What it does
Enable agents to drive a real browser for UI testing, authenticated flows, scraping, and multi-step web automation.
Files
Browser Automation with browser-use CLI
The browser-use command provides fast, persistent browser automation. A background daemon keeps the browser open across commands, giving ~50ms latency per call.
Prerequisites
browser-use doctor # Verify installationFor setup details, see https://github.com/browser-use/browser-use/blob/main/browser_use/skill_cli/README.md
Core Workflow
1. Navigate: browser-use open <url> — launches headless browser and opens page 2. Inspect: browser-use state — returns clickable elements with indices 3. Interact: use indices from state (browser-use click 5, browser-use input 3 "text") 4. Verify: browser-use state or browser-use screenshot to confirm 5. Repeat: browser stays open between commands
If a command fails, run browser-use close first to clear any broken session, then retry.
To use the user's existing Chrome (preserves logins/cookies): run browser-use connect first. To use a cloud browser instead: run browser-use cloud connect first. After either, commands work the same way.
Browser Modes
browser-use open <url> # Default: headless Chromium (no setup needed)
browser-use --headed open <url> # Visible window (for debugging)
browser-use connect # Connect to user's Chrome (preserves logins/cookies)
browser-use cloud connect # Cloud browser (zero-config, requires API key)
browser-use --profile "Default" open <url> # Real Chrome with specific profileAfter connect or cloud connect, all subsequent commands go to that browser — no extra flags needed.
Commands
# Navigation
browser-use open <url> # Navigate to URL
browser-use back # Go back in history
browser-use scroll down # Scroll down (--amount N for pixels)
browser-use scroll up # Scroll up
browser-use tab list # List all tabs
browser-use tab new [url] # Open a new tab (blank or with URL)
browser-use tab switch <index> # Switch to tab by index
browser-use tab close <index> [index...] # Close one or more tabs
# Page State — always run state first to get element indices
browser-use state # URL, title, clickable elements with indices
browser-use screenshot [path.png] # Screenshot (base64 if no path, --full for full page)
# Interactions — use indices from state
browser-use click <index> # Click element by index
browser-use click <x> <y> # Click at pixel coordinates
browser-use type "text" # Type into focused element
browser-use input <index> "text" # Click element, then type
browser-use keys "Enter" # Send keyboard keys (also "Control+a", etc.)
browser-use select <index> "option" # Select dropdown option
browser-use upload <index> <path> # Upload file to file input
browser-use hover <index> # Hover over element
browser-use dblclick <index> # Double-click element
browser-use rightclick <index> # Right-click element
# Data Extraction
browser-use eval "js code" # Execute JavaScript, return result
browser-use get title # Page title
browser-use get html [--selector "h1"] # Page HTML (or scoped to selector)
browser-use get text <index> # Element text content
browser-use get value <index> # Input/textarea value
browser-use get attributes <index> # Element attributes
browser-use get bbox <index> # Bounding box (x, y, width, height)
# Wait
browser-use wait selector "css" # Wait for element (--state visible|hidden|attached|detached, --timeout ms)
browser-use wait text "text" # Wait for text to appear
# Cookies
browser-use cookies get [--url <url>] # Get cookies (optionally filtered)
browser-use cookies set <name> <value> # Set cookie (--domain, --secure, --http-only, --same-site, --expires)
browser-use cookies clear [--url <url>] # Clear cookies
browser-use cookies export <file> # Export to JSON
browser-use cookies import <file> # Import from JSON
# Session
browser-use close # Close browser and stop daemon
browser-use sessions # List active sessions
browser-use close --all # Close all sessionsFor advanced browser control (CDP, device emulation, tab activation), see references/cdp-python.md.
Cloud API
browser-use cloud connect # Provision cloud browser and connect (zero-config)
browser-use cloud login <api-key> # Save API key (or set BROWSER_USE_API_KEY)
browser-use cloud logout # Remove API key
browser-use cloud v2 GET /browsers # REST passthrough (v2 or v3)
browser-use cloud v2 POST /tasks '{"task":"...","url":"..."}'
browser-use cloud v2 poll <task-id> # Poll task until done
browser-use cloud v2 --help # Show API endpointscloud connect provisions a cloud browser with a persistent profile (auto-created on first use), connects via CDP, and prints a live URL. browser-use close disconnects AND stops the cloud browser. For custom browser settings (proxy, timeout, specific profile), use cloud v2 POST /browsers directly with the desired parameters.
Agent Self-Registration
Only use this if you don't already have an API key (check browser-use doctor to see if api_key is set). If already logged in, skip this entirely.
1. browser-use cloud signup — get a challenge 2. Solve the challenge 3. browser-use cloud signup --verify <challenge-id> <answer> — verify and save API key 4. browser-use cloud signup --claim — generate URL for a human to claim the account
Tunnels
browser-use tunnel <port> # Start Cloudflare tunnel (idempotent)
browser-use tunnel list # Show active tunnels
browser-use tunnel stop <port> # Stop tunnel
browser-use tunnel stop --all # Stop all tunnelsProfile Management
browser-use profile list # List detected browsers and profiles
browser-use profile sync --all # Sync profiles to cloud
browser-use profile update # Download/update profile-use binaryCommand Chaining
Commands can be chained with &&. The browser persists via the daemon, so chaining is safe and efficient.
browser-use open https://example.com && browser-use state
browser-use input 5 "user@example.com" && browser-use input 6 "password" && browser-use click 7Chain when you don't need intermediate output. Run separately when you need to parse state to discover indices first.
Common Workflows
Authenticated Browsing
When a task requires an authenticated site (Gmail, GitHub, internal tools), use Chrome profiles:
browser-use profile list # Check available profiles
# Ask the user which profile to use, then:
browser-use --profile "Default" open https://github.com # Already logged inExposing Local Dev Servers
browser-use tunnel 3000 # → https://abc.trycloudflare.com
browser-use open https://abc.trycloudflare.com # Browse the tunnelMultiple Browsers
For subagent workflows or running multiple browsers in parallel, use --session NAME. Each session gets its own browser. See references/multi-session.md.
Configuration
browser-use config list # Show all config values
browser-use config set cloud_connect_proxy jp # Set a value
browser-use config get cloud_connect_proxy # Get a value
browser-use config unset cloud_connect_timeout # Remove a value
browser-use doctor # Shows config + diagnostics
browser-use setup # Interactive post-install setupConfig stored in ~/.browser-use/config.json.
Global Options
| Option | Description |
|---|---|
--headed | Show browser window |
--profile [NAME] | Use real Chrome (bare --profile uses "Default") |
--cdp-url <url> | Connect via CDP URL (http:// or ws://) |
--session NAME | Target a named session (default: "default") |
--json | Output as JSON |
--mcp | Run as MCP server via stdin/stdout |
Tips
1. Always run `state` first to see available elements and their indices 2. Use `--headed` for debugging to see what the browser is doing 3. Sessions persist — browser stays open between commands 4. CLI aliases: bu, browser, and browseruse all work 5. If commands fail, run browser-use close first, then retry
Troubleshooting
- Browser won't start?
browser-use closethenbrowser-use --headed open <url> - Element not found?
browser-use scroll downthenbrowser-use state - Run diagnostics:
browser-use doctor
Cleanup
browser-use close # Close browser session
browser-use tunnel stop --all # Stop tunnels (if any)Raw CDP & Python Session Reference
The CLI commands handle most browser interactions. Use browser-use python with raw CDP when you need browser-level control the CLI doesn't expose — activating a tab so the user sees it, intercepting network requests, emulating devices, or working with Chrome target IDs directly.
How the Python session works
browser-use python "statement" executes one Python statement per call. Variables persist across calls — set a value in one call, use it in the next.
A browser object is pre-injected with sync wrappers for common operations (browser.goto(), browser.click(), etc.). For anything beyond those, two internals give you full access:
browser._run(coroutine)— run any async coroutine synchronously (60s timeout)browser._session— the rawBrowserSessionwith full CDP client access
Getting a CDP client
browser-use python "cdp = browser._run(browser._session.get_or_create_cdp_session())"After this, cdp persists across calls. Use cdp.cdp_client.send.<Domain>.<method>() for any CDP command and cdp.session_id for the session parameter.
Recipes
Activate a tab (make it visible to the user)
The CLI's tab switch only changes the agent's internal focus — Chrome's visible tab doesn't change. To actually show the user a specific tab:
# Get all targets to find the target ID
browser-use python "targets = browser._session.session_manager.get_all_page_targets()"
browser-use python "print([(i, t.url) for i, t in enumerate(targets)])"
# Activate target at index 1 so the user sees it
browser-use python "cdp = browser._run(browser._session.get_or_create_cdp_session(target_id=None, focus=False))"
browser-use python "browser._run(cdp.cdp_client.send.Target.activateTarget(params={'targetId': targets[1].target_id}))"List all tabs with target IDs
browser-use python "targets = browser._session.session_manager.get_all_page_targets()"
browser-use python "
for i, t in enumerate(targets):
print(f'{i}: {t.target_id[:12]}... {t.url}')
"Run JavaScript and get the result
browser-use python "cdp = browser._run(browser._session.get_or_create_cdp_session())"
browser-use python "result = browser._run(cdp.cdp_client.send.Runtime.evaluate(params={'expression': 'document.title', 'returnByValue': True}, session_id=cdp.session_id))"
browser-use python "print(result['result']['value'])"Emulate a mobile device
browser-use python "cdp = browser._run(browser._session.get_or_create_cdp_session())"
browser-use python "browser._run(cdp.cdp_client.send.Emulation.setDeviceMetricsOverride(params={'width': 375, 'height': 812, 'deviceScaleFactor': 3, 'mobile': True}, session_id=cdp.session_id))"Get cookies via CDP
browser-use python "cdp = browser._run(browser._session.get_or_create_cdp_session())"
browser-use python "cookies = browser._run(cdp.cdp_client.send.Network.getCookies(params={}, session_id=cdp.session_id))"
browser-use python "print(cookies)"Tips
- Each
browser-use pythoncall is one statement. Multi-line strings work forforloops andifblocks, but you can't mix statements and expressions. Use multiple calls. - Variables persist: set
cdp = ...in one call, usecdpin the next. - The
browser._run()bridge has a 60-second timeout. For long operations, increase it or use the async internals directly. - All CDP domains are available via
cdp.cdp_client.send.<Domain>.<method>(). See the Chrome DevTools Protocol docs for the full API.
Multiple Browser Sessions
Why use multiple sessions
When you need more than one browser at a time:
- Cloud browser for scraping + local Chrome for authenticated tasks
- Two different Chrome profiles simultaneously
- Isolated browser for testing that won't affect the user's browsing
- Running a headed browser for debugging while headless runs in background
How sessions are isolated
Each --session NAME gets:
- Its own daemon process
- Its own Unix socket (
~/.browser-use/{name}.sock) - Its own PID file and state file
- Its own browser instance (completely independent)
- Its own tab ownership state (multi-agent locks don't cross sessions)
The --session flag
Must be passed on every command targeting that session:
browser-use --session work open <url> # goes to 'work' daemon
browser-use --session work state # reads from 'work' daemon
browser-use state # goes to 'default' daemon (different browser)If you forget --session, the command goes to the default session. This is the most common mistake — you'll interact with the wrong browser.
Combining sessions with browser modes
# Session 1: cloud browser
browser-use --session cloud cloud connect
# Session 2: connect to user's Chrome
browser-use --session chrome connect
# Session 3: headed Chromium for debugging
browser-use --session debug --headed open <url>Each session is fully independent. The cloud session talks to a remote browser, the chrome session talks to the user's Chrome, and the debug session manages its own Chromium — all running simultaneously.
Listing and managing sessions
browser-use sessionsOutput:
SESSION PHASE PID CONFIG
cloud running 12345 cloud
chrome running 12346 cdp
debug ready 12347 headedPHASE shows the daemon lifecycle state: initializing, ready, starting, running, shutting_down, stopped, failed.
browser-use --session cloud close # close one session
browser-use close --all # close every sessionCommon patterns
Cloud + local authenticated:
browser-use --session scraper cloud connect
browser-use --session scraper open https://example.com
# ... scrape data ...
browser-use --session auth --profile "Default" open https://github.com
browser-use --session auth state
# ... interact with authenticated site ...Throwaway test browser:
browser-use --session test --headed open https://localhost:3000
# ... test, debug, inspect ...
browser-use --session test close # done, clean upEnvironment variable:
export BROWSER_USE_SESSION=work
browser-use open <url> # uses 'work' session without --session flag