Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape a JavaScript-driven site with Selenium and Python, run a real browser, wait for the exact DOM state your data needs, extract only the required fields, and always call driver.quit(). A reliable workflow is: create an isolated Python environment, install Selenium, start a WebDriver session, navigate, locate stable elements, synchronize with explicit waits, handle pagination or interactions, checkpoint results, and terminate the browser in a finally block. Selenium is a browser-automation interface implemented through language bindings; WebDriver is a W3C Recommendation. Its WebDriver BiDi work also exposes bidirectional events such as network requests, console messages and JavaScript errors.

Use Selenium only where browser automation is justified

Selenium executes a real browser, so it is useful when JavaScript, client-side rendering, login flows or user interactions produce the content you need. A direct HTTP client is usually simpler for a static page or a documented endpoint. Choose Selenium when browser fidelity matters, and accept the extra startup, memory and synchronization cost.

Before collecting anything, check the target site’s terms, robots guidance, authentication rules and rate limits, together with the law that applies to your use and location. Do not defeat access controls, CAPTCHAs or bot checks. The technical behavior of Selenium does not grant permission to collect a site’s data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Python, Selenium and a browser

Supported versions and drivers

The current Selenium Python API documentation lists Selenium 4.49.0 and Python 3.10 or newer. It lists Chrome, Edge, Firefox, Safari, WebKitGTK and WPEWebKit among supported browsers. Selenium Manager generally obtains a compatible driver when you instantiate a WebDriver, so a separate driver download is often unnecessary; keep the browser and Selenium package current on machines that run unattended jobs.

Create an isolated environment

  1. Install Python 3.10 or newer and a supported browser.
  2. Create and activate a virtual environment: python -m venv .venv, then use .venvScriptsactivate on Windows or source .venv/bin/activate on macOS and Linux.
  3. Install or upgrade the binding: python -m pip install -U selenium.
  4. Verify the installation with python -c "import selenium; print(selenium.__version__)".

A minimal, safe session

from selenium import webdriver
from selenium.webdriver.common.by import By

driver = webdriver.Chrome()
try:
    driver.get('https://example.com')
    heading = driver.find_element(By.TAG_NAME, 'h1').text
    print(heading)
finally:
    driver.quit()

The sequence is deliberately small: import webdriver, create a browser, call get, locate an element, read it, and release the complete session with quit(). Production code should retain the finally block even when extraction grows more complex.

Build a complete scraper around an explicit page state

The following example demonstrates a maintainable pattern. Replace the URL and selectors with ones you are authorized to use. It waits for article cards, normalizes text, writes a checkpoint after each page, and stops when the next link is absent.

import json
from pathlib import Path

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, StaleElementReferenceException

START_URL = 'https://example.com/catalog'
OUT = Path('records.jsonl')
WAIT_SECONDS = 15

options = webdriver.ChromeOptions()
# options.add_argument('--headless=new')  # enable on a server after local testing

driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, WAIT_SECONDS)
seen = set()

try:
    driver.get(START_URL)
    while True:
        cards = wait.until(
            EC.presence_of_all_elements_located(
                (By.CSS_SELECTOR, 'article[data-id]')
            )
        )

        with OUT.open('a', encoding='utf-8') as fh:
            for card in cards:
                key = card.get_attribute('data-id')
                if key in seen:
                    continue
                record = {
                    'id': key,
                    'title': card.find_element(
                        By.CSS_SELECTOR, '[data-role="title"]'
                    ).text.strip(),
                    'url': card.find_element(
                        By.CSS_SELECTOR, 'a[href]'
                    ).get_attribute('href'),
                }
                fh.write(json.dumps(record, ensure_ascii=False) + 'n')
                seen.add(key)

        old_count = len(cards)
        try:
            next_link = driver.find_element(
                By.CSS_SELECTOR, 'a[rel="next"]'
            )
        except Exception:
            break
        if not next_link.is_enabled():
            break

        old_url = driver.current_url
        driver.execute_script('arguments[0].click();', next_link)
        wait.until(lambda d: d.current_url != old_url)
        wait.until(
            lambda d: len(d.find_elements(
                By.CSS_SELECTOR, 'article[data-id]'
            )) != old_count
        )
except (TimeoutException, StaleElementReferenceException) as exc:
    print(f'Page state did not settle: {exc}')
finally:
    driver.quit()

The example uses a stable data-id as a deduplication key and appends each record immediately. If a browser or network failure occurs, already-written lines remain available for a restart. In a real project, catch only the exceptions you can recover from and log the URL, selector and page number with every failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand navigation and page-load state

driver.get(url) waits for the browser’s page-load event before returning. That event is only an initial milestone: JavaScript and AJAX can continue to add or replace DOM nodes afterward. Synchronize on the state that proves your extraction is ready, such as a card count, a result heading, a populated table or the disappearance of a loading indicator.

Choose a page-load strategy deliberately

Strategy When navigation returns What your code must do
normal After the normal page-load sequence. Still wait for application data that arrives through JavaScript or AJAX.
eager Earlier, after the DOM is available without waiting for every subresource. Add explicit waits for every element or state used by extraction.
none As soon as navigation is initiated. Provide all synchronization yourself; this is easiest to misuse.

Faster strategies can reduce idle time, but they return before more of the page is ready. Validate the chosen strategy with the browser and Selenium version used in deployment. Browser options also cover proxies, viewport and other capabilities; Selenium’s default implicit element-location timeout is zero.

Use locators that survive page changes

Prefer semantic selectors

Keep locator definitions separate from extraction logic. Start with By.ID or By.NAME, then use stable CSS selectors that reflect the page’s semantics. A site-provided data-* attribute is usually more durable than a presentation class. Keep selectors in a small module or dictionary so a markup change has one repair point.

  • Good: [data-testid="product-card"], article[data-id] or a stable form name.
  • Fragile: generated class names that change on every build.
  • Very fragile: absolute XPath expressions tied to a complete DOM tree.

After locating an element, read .text or a specific attribute such as href, then normalize whitespace before writing a record. Do not silently treat a missing element as an empty value; record the URL and decide whether the field is optional.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Synchronize with explicit waits

An explicit wait repeatedly evaluates a condition until it succeeds or its timeout expires. Select the condition that matches the next operation rather than increasing a timeout blindly.

Condition Use it when Example
Presence The node must exist in the DOM; it may still be hidden. EC.presence_of_element_located(locator)
Visibility You will read visible text or dimensions. EC.visibility_of_element_located(locator)
Text A known status or result string signals readiness. EC.text_to_be_present_in_element(locator, 'Loaded')
Clickability The next action is a user-like click. EC.element_to_be_clickable(locator)
Custom state Readiness is a count, URL change or application-specific flag. wait.until(lambda d: len(d.find_elements(*locator)) >= 20)
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

wait = WebDriverWait(driver, 15)
card = wait.until(
    EC.visibility_of_element_located(
        (By.CSS_SELECTOR, 'article[data-id]')
    )
)

Do not mix implicit and explicit waits in one session. Selenium warns that the combined timing is unpredictable; its example of a 10-second implicit wait plus a 15-second explicit wait can take about 20 seconds to time out instead of the nominal 15. Keep the implicit timeout at its zero default and make each explicit wait express the state you actually need.

Handle clicks, infinite scroll and pagination

Click only after the control is ready

Wait for clickability, click the control, then wait for a measurable change. Useful signals include a different URL, a larger result count, a new page number or staleness of the old element. Do not sleep for an arbitrary number of seconds and assume the request completed.

load_more = wait.until(
    EC.element_to_be_clickable(
        (By.CSS_SELECTOR, 'button[data-action="load-more"]')
    )
)
old_count = len(driver.find_elements(By.CSS_SELECTOR, 'article[data-id]'))
load_more.click()
wait.until(lambda d: len(d.find_elements(
    By.CSS_SELECTOR, 'article[data-id]'
)) > old_count)

Protect against stale references

Frameworks often replace a node after an AJAX response. A previously stored WebElement can then raise StaleElementReferenceException. Re-locate the element inside the wait or loop, and extract its values before triggering the next DOM update.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Deduplicate and checkpoint

Use a stable URL, site identifier or API-like data attribute as the key. Write records incrementally, keep the last successful page or cursor, and make a restart skip keys already present. This turns a transient browser failure into a partial retry instead of a complete rerun.

Run headless and configure the browser

Develop with a visible browser first; it makes selectors, redirects and consent dialogs inspectable. Enable headless mode only after the workflow is stable. A headless session can still differ in viewport, fonts, GPU behavior and timing, so set a deliberate window size and test the same options used in production.

from selenium import webdriver

options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,1200')
# options.add_argument('--proxy-server=http://proxy.example:8080')

driver = webdriver.Chrome(options=options)

Use one fresh driver per independent job. A finally block calling quit() closes all windows and releases the session; closing only the current tab can leave a driver process behind.

Diagnose common failures

Symptom Likely cause Fix
NoSuchElementException immediately after get() The node is inserted later or the locator is wrong. Inspect the rendered DOM, correct the selector and wait for presence or visibility.
TimeoutException after a long wait The condition never becomes true, the page is blocked, or the selector targets an old state. Log the URL and HTML screenshot, verify the condition manually, and wait for a state change rather than adding time blindly.
Text is empty but the element exists The node is hidden, its text is in a child updated later, or content is rendered outside the selected node. Wait for visibility or expected text and inspect the relevant attribute or child element.
StaleElementReferenceException A framework replaced the node. Locate it again after the update and avoid retaining WebElements across navigation.
Works visibly, fails headless Different viewport, timing, fonts or browser flags. Set a fixed window size, capture diagnostic output, and test headless with the production options.
Driver or browser version error Incompatible browser, binding or driver. Upgrade Selenium and the browser together, or use Selenium Manager’s resolved driver on a clean host.
Browser processes remain after a crash Teardown was skipped. Put quit() in finally; clean orphaned processes before retrying a job.

Know when to use Remote WebDriver or Grid

Remote WebDriver sends commands to a browser session running elsewhere. Selenium Grid provides the routing and execution layer for sessions on remote machines, which is useful for CI isolation, different operating systems or parallel jobs. It is infrastructure, not a requirement for a small local scraper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stay local when

  • One job runs at a time and fits on the development or worker machine.
  • You need the simplest debugging path and can keep a supported browser installed.
  • Your target’s rate limits make high concurrency inappropriate.

Move to Grid when

  • Independent URLs can be processed concurrently without violating site limits.
  • CI workers cannot host a full browser reliably.
  • You need a controlled matrix of browser or operating-system versions.

Grid adds network failure modes, session capacity planning and artifact collection. Start with a bounded worker count, explicit per-page timeouts and retry rules that do not duplicate records.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability and cost trade-offs

A real browser consumes more CPU, memory and startup time than an HTTP client. Reduce waste by extracting only needed fields, avoiding unnecessary navigation, reusing a session for pages in the same authorized workflow, and checkpointing frequently. Do not reuse one driver across unrelated jobs or threads unless your execution design explicitly serializes access.

Reliability comes from deterministic states rather than large delays: wait for the next operation’s condition, record diagnostics on failure, and retry only transient navigation or network errors. A retry should reopen the page and re-check the deduplication key, not blindly append another copy.

For a comparison with a direct HTTP approach, evaluate JavaScript requirements, browser fidelity, startup and resource cost, locator and wait complexity, concurrency and remote execution, debugging visibility, and the target site’s permissions and rate limits. Selenium is strongest when a real browser interaction is necessary; an HTTP client is simpler when a documented, stable endpoint already supplies the data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your goal is a rendered image or PDF rather than structured fields, ScreenshotNeo provides a one-request website screenshot API and an MCP server for AI agents. It is not a replacement for Selenium’s element-level extraction, but it avoids maintaining a browser when you only need a clean visual capture.

cURL (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and whether it was billed. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Sign up free for ScreenshotNeo.

FAQ

Does WebDriver BiDi change how I extract page text?

No. Standard WebDriver commands and Selenium’s Python element APIs still perform navigation and extraction. BiDi adds event streams such as network, console and JavaScript-error notifications, which can improve diagnostics when a page fails without an obvious DOM symptom.

Can a longer timeout fix every flaky scraper?

No. A timeout helps only if the condition eventually becomes true. If the selector is wrong, a request is blocked or the page entered an error state, a longer wait delays the same failure. Identify and wait for the missing state, and capture diagnostics when it is absent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is Grid required to scrape a JavaScript site?

No. A local WebDriver session handles a single job. Grid becomes useful for remote execution, CI isolation or controlled parallelism, with the additional operational cost described above.

Frequently Asked Questions

Does WebDriver BiDi change how I extract page text?

No. Standard WebDriver commands and Selenium’s Python element APIs still perform navigation and extraction. BiDi adds event streams such as network, console and JavaScript-error notifications, which can improve diagnostics when a page fails without an obvious DOM symptom.

Can a longer timeout fix every flaky scraper?

No. A timeout helps only if the condition eventually becomes true. If the selector is wrong, a request is blocked or the page entered an error state, a longer wait delays the same failure. Identify and wait for the missing state, and capture diagnostics when it is absent.

Is Grid required to scrape a JavaScript site?

No. A local WebDriver session handles a single job. Grid becomes useful for remote execution, CI isolation or controlled parallelism, with additional operational cost.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.