Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Yes—Selenium can scrape JavaScript-rendered sites with Python. It launches a real browser, waits for the application state you need, and then reads the same DOM a user sees. This guide builds a small scraper that collects article titles and links, then expands it with robust locators, explicit waits, pagination, sessions, retries, checkpoints, and responsible-use safeguards.

What you will build

Our example collects article cards from a page whose content is rendered or updated by JavaScript. The pattern applies to product listings, dashboards, search results, and other sites where a plain HTTP request returns incomplete HTML.

  • Python 3.10 or newer.
  • A current installation of Chrome, Edge, Firefox, Safari, WebKitGTK, or WPEWebKit.
  • A virtual environment and the Selenium Python package.

Current Selenium Python documentation supports Python 3.10+ and Selenium 4.49.0 documentation describes browser automation across those browser families. Modern Selenium includes Selenium Manager, which usually obtains a compatible driver automatically.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

1. Create an isolated project

mkdir selenium-scraper
cd selenium-scraper
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install -U selenium

Keep your scraper and its output in this directory. A lock file or pinned package version is sensible for production, but update Selenium regularly so browser compatibility fixes are included.

2. Launch a browser and inspect the page

Start with webdriver.Chrome(). Selenium Manager handles driver discovery in common cases, so a separate ChromeDriver download is not normally required.

from selenium import webdriver

browser = webdriver.Chrome()
try:
    browser.get("https://example.com")
    print(browser.title)
    print(browser.current_url)
finally:
    browser.quit()

Replace the URL with a page you are permitted to access. Inspect it in your browser’s developer tools: right-click a target, choose Inspect, and identify the smallest stable element that contains the data. Examine the live DOM after JavaScript has run, not only the original page source.

3. Choose locators that survive redesigns

Prefer a stable ID

Selenium’s locator guidance says: “In general, if HTML IDs are available, unique, and consistently predictable, they are the preferred method for locating an element on a page.” An ID such as article-list is usually clearer and less fragile than a long class chain.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use compact CSS next

When no reliable ID exists, use a short CSS selector such as article.card h2 a. Avoid presentation-only classes, generated names, and selectors that depend on an exact nesting depth.

Use XPath for relationships

XPath is useful when you must locate an element by relationship or text, but Selenium describes it as harder to debug and typically slower. Keep it narrow and readable, for example //article[.//h2]/descendant::a[1].

Locator Best use Main risk
Unique ID Stable, direct lookup Not present or regenerated
CSS selector Readable combinations and descendants Breaks when classes are renamed
XPath Text and structural relationships Harder maintenance and debugging

4. Wait for the state you actually need

driver.get() returning only tells you that the selected page-load policy has completed; it does not prove that an XHR, fetch call, click, infinite-scroll operation, or client-side route has populated your target.

Use an explicit wait

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

wait = WebDriverWait(browser, 10)
list_element = wait.until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "article"))
)
wait.until(EC.visibility_of(list_element))

Presence means the node exists in the DOM; visibility also requires it to be displayed. For dynamic content, add a condition that proves the right data arrived:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
wait.until(
    EC.text_to_be_present_in_element(
        (By.CSS_SELECTOR, "article h2"), "Python"
    )
)

Prefer condition-based synchronization to arbitrary time.sleep(). A fixed sleep either under-waits on a slow run or wastes time on a fast one.

Do not combine wait policies accidentally

Selenium’s waiting guidance states exactly: “Do not mix implicit and explicit waits.” An implicit wait changes how every element lookup polls, making explicit timeout behavior difficult to predict. Choose explicit waits for a scraper whose states differ by page or component.

5. A complete scraper

This script waits for cards, extracts text and links, follows a “Next” control until it disappears, retries transient page failures, and checkpoints each page to JSON Lines. Replace the selectors with those from your target DOM.

import json
import time
from pathlib import Path
from urllib.parse import urljoin

from selenium import webdriver
from selenium.common.exceptions import TimeoutException, WebDriverException
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

START_URL = "https://example.com/news"
CARD = "article.card"
NEXT = "a[rel='next']"
OUTPUT = Path("articles.jsonl")

options = webdriver.ChromeOptions()
# options.add_argument("--headless=new")  # enable for unattended runs
options.page_load_strategy = "normal"  # normal, eager, or none
browser = webdriver.Chrome(options=options)
browser.set_page_load_timeout(45)
browser.set_script_timeout(30)
wait = WebDriverWait(browser, 15)


def read_cards():
    wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, CARD)))
    rows = []
    for card in browser.find_elements(By.CSS_SELECTOR, CARD):
        heading = card.find_element(By.CSS_SELECTOR, "h2")
        link = heading.find_element(By.CSS_SELECTOR, "a")
        href = link.get_attribute("href")
        rows.append({
            "title": heading.text.strip(),
            "url": urljoin(browser.current_url, href),
        })
    return rows


def next_page():
    try:
        control = browser.find_element(By.CSS_SELECTOR, NEXT)
        if not control.is_displayed() or not control.is_enabled():
            return False
        old_url = browser.current_url
        browser.execute_script("arguments[0].click();", control)
        wait.until(lambda d: d.current_url != old_url or
                   len(d.find_elements(By.CSS_SELECTOR, CARD)) > 0)
        return True
    except (TimeoutException, WebDriverException):
        return False


try:
    browser.get(START_URL)
    seen = set()
    with OUTPUT.open("w", encoding="utf-8") as stream:
        for page_number in range(1, 101):
            for attempt in range(3):
                try:
                    records = read_cards()
                    break
                except (TimeoutException, WebDriverException):
                    if attempt == 2:
                        raise
                    time.sleep(2 ** attempt)
            new_records = [r for r in records if r["url"] not in seen]
            for record in new_records:
                seen.add(record["url"])
                record["page"] = page_number
                stream.write(json.dumps(record, ensure_ascii=False) + "n")
            stream.flush()  # checkpoint survives an interrupted run
            if not next_page():
                break
finally:
    browser.quit()

6. Extract more than visible text

Attributes and links

Use get_attribute() for URLs, image sources, data attributes, and ARIA labels. Resolve relative links with urljoin(), as the example does.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tables

for row in browser.find_elements(By.CSS_SELECTOR, "table tbody tr"):
    cells = [cell.text.strip() for cell in
             row.find_elements(By.CSS_SELECTOR, "th, td")]
    print(cells)

Shadow DOM and iframes

For an iframe, switch first with browser.switch_to.frame(frame_element), scrape its document, then call browser.switch_to.default_content(). Components in an open shadow root require Selenium’s shadow-root APIs; closed roots are not directly queryable from page JavaScript.

7. Page-load strategy, timeouts, and browser options

Strategy Browser returns after Use when
normal The load event and dependent resources finish Safest default
eager DOMContentLoaded Images and nonessential assets are irrelevant
none Navigation starts without blocking You have strong, explicit readiness checks

Set page-load, script, and element timeouts deliberately. A short page-load timeout does not replace an explicit wait for the card that appears after an API call. A proxy can be configured through browser options when a restricted network, traffic capture, or mock backend requires it.

8. Pagination, sessions, and polite throughput

  • Preserve the same driver so cookies, login state, and local storage remain available.
  • Stop on a missing, disabled, or unchanged “Next” control; also deduplicate URLs to prevent loops.
  • Retry transient navigation failures with a small capped backoff, as in the example. Do not retry indefinitely.
  • Flush checkpoints after each page so a crash does not discard earlier work.
  • For infinite scroll, scroll in measured increments and wait for the item count to increase before continuing.
  • Limit concurrency and add delays appropriate to the site. A real browser consumes substantially more CPU, memory, and bandwidth than a static HTTP client.

9. Diagnose common failures

“NoSuchElementException”

The selector may be wrong, the element may be inside an iframe or shadow root, or the app has not rendered it. Inspect the post-JavaScript DOM, switch into the correct frame, and wait for a meaningful condition before lookup.

“Element not interactable” or click interception

A modal, sticky header, animation, or consent layer may cover the control. Wait for clickability, scroll it into view, close the overlay when permitted, and verify that the click changed the expected state. JavaScript clicking should be a last resort because it can bypass normal user-event behavior.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TimeoutException

Check whether the selector is stale, the request failed, a bot check appeared, or the page uses a different route than expected. Capture a screenshot and page source on failure, then increase a timeout only after identifying the slow state.

Driver or browser version mismatch

Update the browser and Selenium package, allow Selenium Manager to resolve the driver, and remove an obsolete driver executable from your PATH. Manual driver installation is still appropriate in locked-down environments; match the browser’s major version and point the service explicitly.

Blank or partial data

Confirm that your wait checks the content rather than merely document.readyState. Increase the script timeout for long client operations, check network access, and verify that lazy-loaded items were actually scrolled into view.

10. Responsible scraping

Read the site’s terms and access rules before automating. Inspect robots.txt; RFC 9309 is the IETF reference for the Robots Exclusion Protocol. Robots rules are an access signal, not a blanket legal determination, so obtain permission where required and stop when a site blocks automation. Identify your user agent where appropriate, respect rate limits, and avoid collecting personal data you do not need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you only need a clean image or PDF rather than DOM-level extraction, ScreenshotNeo provides a single-call website screenshot API. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all options. The same endpoint supports full-page and element captures, device and viewport settings, retina scale, PDFs, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data, and an OpenAPI specification. Parameters commonly used by other screenshot APIs also work, easing migration.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account to try it.

FAQ

Do I still need ChromeDriver?

Usually not: Selenium Manager handles common driver installation. Manual drivers remain useful when policy, networking, or a pinned browser image prevents automatic resolution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which wait should I choose for a single scraper?

Use explicit waits tied to the state you need and avoid mixing them with implicit waits. This keeps each page’s synchronization visible and predictable.

When is a static HTTP client better?

Use one when the needed data is present in the initial HTML or a documented API. It is lighter and faster; choose Selenium when JavaScript execution, user interactions, or browser session behavior is essential.

Frequently Asked Questions

Can Selenium scrape content loaded only after scrolling?

Yes. Scroll incrementally, then explicitly wait for the item count or a sentinel element to change before extracting the newly loaded nodes.

How should I make a scraper restartable?

Write each completed record or page to a checkpoint file, deduplicate by a stable URL or ID, and resume from the last confirmed page instead of rebuilding the entire run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.