Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
brightdata avatar

Bright Data Best Practices

  • 1.6k installs
  • 240 repo stars
  • Updated June 25, 2026
  • brightdata/skills

bright-data-best-practices guides production Bright Data API selection, authentication, and scraping patterns for assistants.

About

The bright-data-best-practices skill is reference documentation for coding assistants implementing Bright Data at scale. It maps use cases to four APIs: Web Unlocker for HTTP page fetch with bot bypass, SERP API for Google/Bing/Yandex results, Web Scraper API for pre-built Amazon/LinkedIn/Instagram/TikTok datasets, and Browser API for click/scroll/form automation via Puppeteer, Playwright, or Selenium CDP. Shared auth uses BRIGHTDATA_API_KEY, zone names, and BROWSER_AUTH; the bdata CLI login path is canonical per references/cli-setup.md. Guidance covers choosing the most specific API, REST Bearer headers, proxy network fallbacks, and when to delegate to the proxy.md skill for raw DC/ISP/residential/mobile routing. Agents consult CLI setup before shelling out to bdata and pick Unlocker over Browser API when no interaction is needed. Use when users build scrapers, unblock bot detection, automate SERP collection, or wire Bright Data into Claude Code or Cursor projects.

  • API selection matrix: Unlocker, SERP, Web Scraper, and Browser APIs by interaction depth.
  • Shared authentication env vars and bdata CLI login as canonical setup path.
  • Web Unlocker for cheapest HTTP scrape; Browser API for JS, forms, and XHR intercept.
  • Structured marketplace scrapers without custom parsing via Web Scraper API.
  • CLI setup reference required before any bdata shell invocation.

Bright Data Best Practices by the numbers

  • 1,628 all-time installs (skills.sh)
  • +11 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #306 of 4,347 Backend & APIs skills by installs in the Skillselion catalog
  • Security screen: HIGH risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

bright-data-best-practices capabilities & compatibility

Capabilities
api selection guidance · authentication setup · cli and rest patterns · bot bypass scraping · structured marketplace extraction
Works with
chrome
Use cases
web scraping · research · api development
npx skills add https://github.com/brightdata/skills --skill bright-data-best-practices

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.6k
repo stars240
Security audit2 / 3 scanners passed
Last updatedJune 25, 2026
Repositorybrightdata/skills

Which Bright Data API and auth pattern should I use for scraping, SERP, or browser automation?

Implement production-ready Bright Data integrations for web scraping, SERP extraction, structured marketplace data, and browser automation with correct API selection and auth.

Who is it for?

Developers adding Bright Data scraping or automation to agents and backend services.

Skip if: Raw proxy-only routing without managed APIs (use the proxy.md skill instead).

When should I use this skill?

User mentions Bright Data, Web Unlocker, SERP API, Scraping Browser, or bdata CLI setup.

What you get

Correct API choice, env configuration, and integration pattern for the target extraction job.

  • cdp connection configuration
  • session rule setup
  • captcha handling pattern

By the numbers

  • Reference covers 15 documented sections including CAPTCHA handling, geolocation, and error codes

Files

SKILL.mdMarkdownGitHub ↗

CLI Setup Reference

Install, authentication, and troubleshooting for the Bright Data CLI (bdata) are documented in a single canonical place:

`references/cli-setup.md`

Consult it before any task that shells out to bdata.

Bright Data APIs

Bright Data provides infrastructure for web data extraction at scale. Four primary APIs cover different use cases — always pick the most specific tool for the job.

Choosing the Right API

Use CaseAPIWhy
Scrape any webpage by URL (no interaction)Web UnlockerHTTP-based, auto-bypasses bot detection, cheapest
Google / Bing / Yandex search resultsSERP APISpecialized for SERP extraction, returns structured data
Structured data from Amazon, LinkedIn, Instagram, TikTok, etc.Web Scraper APIPre-built scrapers, no parsing needed
Click, scroll, fill forms, run JS, intercept XHRBrowser APIFull browser automation
Puppeteer / Playwright / Selenium automationBrowser APIConnects via CDP/WebDriver
Route your own HTTP client through a raw proxy (DC/ISP/Residential/Mobile)Proxy networksWhen you need direct proxy access with your own request logic instead of a managed API — see the proxy.md skill

Authentication Pattern (All APIs)

All APIs share the same authentication model. The env vars below apply to direct REST API integrations — if you are using the bdata CLI, bdata login handles all of these automatically (see `references/cli-setup.md`).

export BRIGHTDATA_API_KEY="your-api-key"         # From Control Panel > Account Settings
export BRIGHTDATA_UNLOCKER_ZONE="zone-name"       # Web Unlocker zone name
export BRIGHTDATA_SERP_ZONE="serp-zone-name"      # SERP API zone name
export BROWSER_AUTH="brd-customer-ID-zone-NAME:PASSWORD"  # Browser API credentials

REST API authentication header for Web Unlocker and SERP API:

Authorization: Bearer YOUR_API_KEY

---

Web Unlocker API

HTTP-based scraping proxy. Best for simple page fetches without browser interaction.

Endpoint: POST https://api.brightdata.com/request

import requests

response = requests.post(
    "https://api.brightdata.com/request",
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={
        "zone": "YOUR_ZONE_NAME",
        "url": "https://example.com/product/123",
        "format": "raw"
    }
)
html = response.text

Key Parameters

ParameterTypeDescription
zonestringZone name (required)
urlstringTarget URL with http:// or https:// (required)
formatstring"raw" (HTML) or "json" (structured wrapper) (required)
methodstringHTTP verb, default "GET"
countrystring2-letter ISO for geo-targeting (e.g., "us", "de")
data_formatstringTransform: "markdown" or "screenshot"
asyncbooleantrue for async mode

Quick Patterns

# Get markdown (best for LLM input)
response = requests.post(
    "https://api.brightdata.com/request",
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={"zone": ZONE, "url": url, "format": "raw", "data_format": "markdown"}
)

# Geo-targeted request
json={"zone": ZONE, "url": url, "format": "raw", "country": "de"}

# Screenshot for debugging
json={"zone": ZONE, "url": url, "format": "raw", "data_format": "screenshot"}

# Async for bulk processing
json={"zone": ZONE, "url": url, "format": "raw", "async": True}

Critical rule: Never use Web Unlocker with Puppeteer, Playwright, Selenium, or anti-detect browsers. Use Browser API instead.

See [references/web-unlocker.md](references/web-unlocker.md) for complete reference including proxy interface, special headers, async flow, features, and billing.

---

SERP API

Structured search engine result extraction for Google, Bing, Yandex, DuckDuckGo.

Endpoint: POST https://api.brightdata.com/request (same as Web Unlocker)

response = requests.post(
    "https://api.brightdata.com/request",
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={
        "zone": "YOUR_SERP_ZONE",
        "url": "https://www.google.com/search?q=python+web+scraping&brd_json=1&gl=us&hl=en",
        "format": "raw"
    }
)
data = response.json()
for result in data.get("organic", []):
    print(result["rank"], result["title"], result["link"])

Essential Google URL Parameters

ParameterDescriptionExample
qSearch queryq=python+web+scraping
brd_jsonParsed JSON outputbrd_json=1 (always use for data pipelines)
glCountry for searchgl=us
hlLanguagehl=en
startPagination offsetstart=10 (page 2), start=20 (page 3)
tbmSearch typetbm=nws (news), tbm=isch (images), tbm=vid (videos)
brd_mobileDevicebrd_mobile=1 (mobile), brd_mobile=ios
brd_browserBrowserbrd_browser=chrome
brd_ai_overviewTrigger AI Overviewbrd_ai_overview=2
uuleEncoded geo locationfor precise location targeting

Note: num parameter is deprecated as of September 2025. Use start for pagination.

Parsed JSON Response Structure

{
  "organic": [{"rank": 1, "global_rank": 1, "title": "...", "link": "...", "description": "..."}],
  "paid": [],
  "people_also_ask": [],
  "knowledge_graph": {},
  "related_searches": [],
  "general": {"results_cnt": 1240000000, "query": "..."}
}

Bing Key Parameters

ParameterDescription
qSearch query
setLangLanguage (prefer 4-letter: en-US)
ccCountry code
firstPagination (increment by 10: 1, 11, 21...)
safesearchoff, moderate, strict
brd_mobileDevice type

Async for Bulk SERP

# Submit
response = requests.post(
    "https://api.brightdata.com/request",
    params={"async": "1"},
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={"zone": SERP_ZONE, "url": "https://www.google.com/search?q=test&brd_json=1", "format": "raw"}
)
response_id = response.headers.get("x-response-id")

# Retrieve (retrieve calls are NOT billed)
result = requests.get(
    "https://api.brightdata.com/serp/get_result",
    params={"response_id": response_id},
    headers={"Authorization": f"Bearer {API_KEY}"}
)

Billing: Pay per 1,000 successful requests only. Async retrieve calls are not billed.

See [references/serp-api.md](references/serp-api.md) for complete reference including Maps, Trends, Reviews, Lens, Hotels, Flights parameters.

---

Web Scraper API

Pre-built scrapers for structured data extraction from 100+ platforms. No parsing logic needed.

Sync Endpoint: POST https://api.brightdata.com/datasets/v3/scrape Async Endpoint: POST https://api.brightdata.com/datasets/v3/trigger

# Sync (up to 20 URLs, returns immediately)
response = requests.post(
    "https://api.brightdata.com/datasets/v3/scrape",
    params={"dataset_id": "YOUR_DATASET_ID", "format": "json"},
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={"input": [{"url": "https://www.amazon.com/dp/B09X7M8TBQ"}]}
)

if response.status_code == 200:
    data = response.json()  # Results ready
elif response.status_code == 202:
    snapshot_id = response.json()["snapshot_id"]  # Poll for completion

Parameters

ParameterTypeDescription
dataset_idstringScraper identifier from the Scraper Library (required)
formatstringjson (default), ndjson, jsonl, csv
custom_output_fieldsstringPipe-separated fields: `url\
include_errorsbooleanInclude error info in results

Request Body

{
  "input": [
    { "url": "https://www.amazon.com/dp/B09X7M8TBQ" },
    { "url": "https://www.amazon.com/dp/B0B7CTCPKN" }
  ]
}

Poll for Async Results

import time

# Trigger
snapshot_id = requests.post(
    "https://api.brightdata.com/datasets/v3/trigger",
    params={"dataset_id": DATASET_ID, "format": "json"},
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={"input": [{"url": u} for u in urls]}
).json()["snapshot_id"]

# Poll
while True:
    status = requests.get(
        f"https://api.brightdata.com/datasets/v3/progress/{snapshot_id}",
        headers={"Authorization": f"Bearer {API_KEY}"}
    ).json()["status"]

    if status == "ready": break
    if status == "failed": raise Exception("Job failed")
    time.sleep(10)

# Download
data = requests.get(
    f"https://api.brightdata.com/datasets/v3/snapshot/{snapshot_id}",
    params={"format": "json"},
    headers={"Authorization": f"Bearer {API_KEY}"}
).json()

Progress status values: startingrunningready | failed Data retention: 30 days. Billing: Per delivered record. Invalid input URLs that fail are still billable.

See [references/web-scraper-api.md](references/web-scraper-api.md) for complete reference including scraper types, output formats, delivery options, and billing details.

---

Browser API (Scraping Browser)

Full browser automation via CDP/WebDriver. Handles CAPTCHA, fingerprinting, and anti-bot detection automatically.

Connection:

  • Playwright/Puppeteer: wss://${AUTH}@brd.superproxy.io:9222
  • Selenium: https://${AUTH}@brd.superproxy.io:9515
const { chromium } = require("playwright-core");

const AUTH = process.env.BROWSER_AUTH;
const browser = await chromium.connectOverCDP(`wss://${AUTH}@brd.superproxy.io:9222`);
const page = await browser.newPage();
page.setDefaultNavigationTimeout(120000); // Always set to 2 minutes

await page.goto("https://example.com", { waitUntil: "domcontentloaded" });
const html = await page.content();
await browser.close();
from playwright.async_api import async_playwright

async with async_playwright() as p:
    browser = await p.chromium.connect_over_cdp(f"wss://{AUTH}@brd.superproxy.io:9222")
    page = await browser.new_page()
    page.set_default_navigation_timeout(120000)
    await page.goto("https://example.com", wait_until="domcontentloaded")
    html = await page.content()
    await browser.close()

Custom CDP Functions

FunctionPurpose
Captcha.solveManually trigger CAPTCHA solving
Captcha.setAutoSolveEnable/disable auto CAPTCHA solving
Proxy.setLocationSet precise geo location (call BEFORE goto)
Proxy.useSessionMaintain same IP across sessions
Emulation.setDeviceApply device profile (iPhone 14, etc.)
Emulation.getSupportedDevicesList available device profiles
Unblocker.enableAdBlockBlock ads to save bandwidth
Unblocker.disableAdBlockRe-enable ads
Input.typeFast text input for bulk form filling
Browser.addCertificateInstall client SSL cert for session
Page.inspectGet DevTools debug URL for live session
// CDP session pattern for custom functions
const client = await page.target().createCDPSession();

// CAPTCHA solve with timeout
const result = await client.send("Captcha.solve", { timeout: 30000 });

// Precise geo location (must be before goto)
await client.send("Proxy.setLocation", {
  latitude: 37.7749,
  longitude: -122.4194,
  distance: 10,
  strict: true
});

// Block unnecessary resources
await client.send("Network.setBlockedURLs", { urls: ["*google-analytics*", "*.ads.*"] });

// Device emulation
await client.send("Emulation.setDevice", { deviceName: "iPhone 14" });

Session Rules

  • One initial navigation per session — new URL = new session
  • Idle timeout: 5 minutes
  • Max duration: 30 minutes

Geolocation

  • Country-level: append -country-us to credentials username
  • EU-wide: append -country-eu (routes through 29+ European countries)
  • Precise: use Proxy.setLocation CDP command (before navigation)

Error Codes

CodeIssueFix
407Wrong portPlaywright/Puppeteer → 9222, Selenium → 9515
403Bad authCheck credentials format and zone type
503Service scalingWait 1 minute, reconnect

Billing: Traffic-based only. Block images/CSS/fonts to reduce costs.

See [references/browser-api.md](references/browser-api.md) for complete reference including all CDP functions, bandwidth optimization, CAPTCHA patterns, and debugging.

---

Detailed References

  • [references/web-unlocker.md](references/web-unlocker.md) — Web Unlocker: full parameter list, proxy interface, special headers, async flow, features, billing, anti-patterns
  • [references/serp-api.md](references/serp-api.md) — SERP API: all Google params (Maps, Trends, Reviews, Lens, Hotels, Flights), Bing params, parsed JSON structure, async, billing
  • [references/web-scraper-api.md](references/web-scraper-api.md) — Web Scraper API: sync vs async, all parameters, polling, scraper types, output formats, billing
  • [references/browser-api.md](references/browser-api.md) — Browser API: connection strings, session rules, all CDP functions, geo-targeting, bandwidth optimization, CAPTCHA, debugging, error codes

Related Skills

  • `brightdata-proxy` — For routing requests through Bright Data's raw proxy networks (Datacenter, ISP, Residential, Mobile) with your own HTTP client instead of a managed API. Covers network/IP-pool selection, the brd-customer-... username format, targeting & sticky-session params, SSL CA setup for Residential/Mobile, and integrations for cURL, Python (requests/httpx/aiohttp/Scrapy), Node (fetch/axios), Playwright, Puppeteer, and Selenium. Hand off to it whenever the task is raw proxy access rather than Web Unlocker / SERP / Web Scraper / Browser API. Escalation order when proxies hit consistent blocks: raw proxy → Web Unlocker → Browser API → Web Scraper API.

Related skills

Forks & variants (1)

Bright Data Best Practices has 1 known copy in the catalog totaling 7 installs. They canonicalize to this original listing.

How it compares

Use for Bright Data proxy-backed browser sessions; use generic Playwright skills when no residential proxy or anti-bot bypass is required.

FAQ

When should I use Web Unlocker vs Browser API?

Use Web Unlocker for simple HTTP page fetch; Browser API when you need clicks, scroll, forms, or JS execution.

How is authentication configured?

Set BRIGHTDATA_API_KEY and zone env vars, or run bdata login for CLI workflows per cli-setup.md.

Which API returns structured Amazon or LinkedIn data?

Web Scraper API provides pre-built marketplace scrapers without custom HTML parsing.

Is Bright Data Best Practices safe to install?

skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

Backend & APIsintegrationsbackend

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.