Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Selenium can scrape pages whose data appears only after JavaScript runs or after a browser interaction. The key is to wait for the data you need—not merely for navigation to finish—and to use selectors that can survive ordinary page changes. This guide explains how to choose waits and locators, handle common failures, and decide when Selenium is more browser automation than your task requires.
What is Selenium, and when should you use it for scraping?
Selenium is an open-source browser-automation suite. Its WebDriver interface lets a program control a browser: navigate to a page, locate elements, click or type, and read the resulting page content. Selenium Grid can distribute browser execution across machines, which is useful for parallel work and CI/CD workflows. Selenium’s overview lists Java, Python, C#, JavaScript, Ruby, and Kotlin among its supported languages.
Selenium is a good fit when the data you need is rendered or changed by JavaScript, or when reaching it requires browser interaction. A plain HTTP request retrieves a response; it does not, by itself, run the page’s client-side application or click its controls. If the information is already present in static HTML and no interaction is needed, a direct HTTP client and HTML parser are usually a lighter starting point.
Free tools Windows power users keep installed
One-click scans. No signup required.
| Target or requirement | Practical starting point |
|---|---|
| Data is in the initial HTML; no browser interaction is needed | Direct HTTP request and HTML parser |
| Data appears after JavaScript runs | Selenium, with a wait for the actual data condition |
| Content requires clicking, typing, or another browser action | Selenium WebDriver |
| Many browser tasks must run in parallel or in CI | Consider Selenium Grid or a managed cloud Grid |
Using a browser adds execution and synchronization complexity. Before choosing it, inspect whether the target’s data is available without rendering and whether a supported, less resource-intensive access method meets your needs.
#1 Best Overall
Why does Selenium say the page loaded when the data is missing?
A completed navigation is not a promise that a modern page has finished rendering its useful content. Selenium’s page-load behavior is tied to browser navigation milestones such as the document ready state or load event. JavaScript can continue to fetch data and modify the DOM after those milestones, especially in single-page applications.
The Selenium documentation’s Waiting Strategies page puts it this way: “The readyState only concerns itself with loading assets defined in the HTML, but loaded JavaScript assets often result in changes to the site, and elements that need to be interacted with may not yet be on the page when the code is ready to execute the next Selenium command.”
Instead of treating “page loaded” as “scrape ready,” define the next operation’s success condition. Depending on the page, that might be a result element appearing, a loading indicator disappearing, or a result count changing. A wait for the data is more meaningful than a wait for an arbitrary number of seconds.
Recommended Free Tools
Should you use sleep, implicit waits, or explicit waits?
Explicit waits: the usual choice
An explicit wait polls a chosen condition and proceeds when that condition becomes true or the timeout is reached. It keeps the wait connected to the state your code actually needs. In Python, Selenium’s WebDriverWait API documents a default poll_frequency of 0.5 seconds; until() repeatedly calls the supplied function until it returns a truthy result or times out.
Implicit waits: a global lookup setting
An implicit wait applies to element lookups throughout the driver session. Because it changes the behavior of lookups globally, it can make timing harder to reason about when different elements have different readiness conditions.
Fixed sleeps: a blunt fallback
time.sleep() pauses for a fixed duration whether the page becomes ready quickly or slowly. If the pause is too short, the script races ahead; if it is too long, every run wastes time. A short, deliberate pause may occasionally be useful for a known animation or a temporary diagnostic, but it should not replace condition-based synchronization.
Selenium explicitly warns: “Do not mix implicit and explicit waits. Doing so can cause unpredictable wait times.” Choose one synchronization approach for element readiness; for most scrapers, explicit waits make the intended condition clearest.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →How do you scrape JavaScript-rendered content with Python?
The following example opens a page, waits until article cards exist, and reads their text. Replace the example URL and CSS selector with ones observed on your target. It deliberately waits for a page-specific condition rather than assuming that navigation completion means the data is ready.
Rank #3
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URL = "https://example.com/news"
CARD_SELECTOR = "article[data-test='story-card']"
options = webdriver.ChromeOptions()
# Uncomment for a browser without a visible window, if supported in your setup.
# options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
driver.get(URL)
cards = WebDriverWait(driver, 20).until(
EC.presence_of_all_elements_located((By.CSS_SELECTOR, CARD_SELECTOR))
)
for card in cards:
print(card.text)
finally:
driver.quit()
This is a template, not a universal selector: the example attribute may not exist on the site you want to scrape. Choose a condition that signals the usable state. Presence means an element is in the DOM; it does not necessarily mean it is visible or contains final data. If you need a user-visible result, wait for visibility instead. If the page initially shows a loading state, wait for the relevant result to appear or the loader to disappear.
For a multi-step flow, synchronize each transition separately: wait for the control you need, act on it, then wait for the resulting state before extracting. This makes failures easier to diagnose than a single long pause at the beginning.
Which Selenium locator should you use?
Prefer a unique, stable ID when the page provides one. Selenium’s locator guidance says: “In general, if HTML IDs are available, unique, and consistently predictable, they are the preferred method for locating an element on a page.” When there is no suitable ID, a compact CSS selector based on stable attributes such as data-test or name is often readable and maintainable.
XPath is useful when the relationship between elements or a text condition matters. Keep it narrow enough that another developer can understand what it selects. Avoid absolute paths tied to the entire DOM structure and generated class names that may change between builds. Selenium’s guidance describes XPath as more complicated and typically slower than CSS, so use it when its expressive power helps rather than by default.
- Good candidate: a unique, predictable ID.
- Also useful: a short CSS selector using a stable semantic attribute.
- Use when needed: a focused XPath for a genuine relationship or text-based condition.
- Fragile choices: long absolute XPath expressions and transient generated classes.
When a locator stops working, first inspect the current page structure and confirm whether the element is in the main document or whether the page state has changed. Then update the smallest possible part of the selector. Do not silently broaden a selector until it matches: that can cause the scraper to collect the wrong element without an obvious error.
What do normal, eager, and none page-load strategies change?
Selenium’s page-load strategies determine when navigation returns; they do not identify when a particular piece of scraped data is ready.
| Strategy | Navigation waits for | What your scraper still needs |
|---|---|---|
normal (default) |
The load event / complete ready state | A condition for JavaScript-rendered results or later changes |
eager |
DOMContentLoaded | Explicit synchronization for the content your task depends on |
none |
No document-loading wait | A deliberate wait before interacting with or reading the page |
Using eager or none can avoid waiting for resources that do not matter to the task, such as images. The trade-off is that navigation may return before the DOM or data you need is ready, so your explicit conditions become essential. Keep the default if its navigation behavior is appropriate; do not change strategy as a substitute for understanding the page’s actual readiness signal.
Why is a Selenium click intercepted or not interactable?
A click failure often means the target is not currently usable in the same way a person could use it. It may be hidden, outside the viewport, covered by an overlay, or not yet exposed to pointer or keyboard interaction. Selenium checks whether an element is displayed and interactable, and it can scroll an element into view; that does not make a hidden or covered control clickable.
Best Value
- Wait for the target’s relevant visibility or interactability condition, not just its presence in the DOM.
- Check whether a modal, consent banner, loading layer, or other overlay is covering it; wait for that overlay to disappear or handle it through the site’s normal interface.
- Confirm that the control is the intended one and is enabled in the current page state.
- After clicking, wait for the resulting state before continuing. Do not treat a successful click command as proof that the page completed the action.
Trying to force a click with a different mechanism can bypass the user-facing behavior your scraper is meant to observe and can hide the real problem. First establish why the element is unavailable.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should you choose an approach, and what should you check before scraping?
Make the choice around the target, not the popularity of a tool. Selenium is strongest when a real browser is required; it is unnecessary overhead when the needed content is already in static HTML. Consider four questions:
- Rendering: Is the content in the initial HTML, or does JavaScript create it later?
- Interaction: Must the workflow click, type, authenticate, or change page state?
- Synchronization: Can you identify a stable selector and an observable ready condition?
- Scale: Is one browser session enough, or do you need parallel execution through Grid or a managed cloud Grid?
Before running automation, check the target site’s terms, robots policy, rate limits, authentication requirements, privacy obligations, and applicable law. Whether scraping is permitted depends on the particular site and jurisdiction; Selenium itself does not answer that question. Follow access rules, avoid unnecessary request volume, and do not treat a public page as blanket permission to collect or reuse its contents.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Troubleshooting common Selenium scraping failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Navigation returns, but expected data is absent | JavaScript is still changing the DOM after the navigation milestone | Wait for the result element, changed count, or loader state that represents usable data. |
| Element lookup times out | Wrong selector, wrong page state, or content not yet present | Inspect the current DOM and selector; add a condition-based wait for the expected state. |
| Locator matches the wrong item or multiple items | Selector is too broad or depends on unstable markup | Use a unique stable ID or narrow CSS selector; verify the selected element before extracting. |
| Click intercepted or element not interactable | Hidden/off-screen control, overlay, or unavailable interaction state | Wait for visibility/interactability, resolve the overlay, and verify the control is enabled. |
| Runs take unpredictably long | Fixed sleeps or mixed implicit and explicit waits | Use explicit waits for named conditions and avoid mixing wait types. |
| Scrape works locally but struggles at larger scale | Each task requires browser execution and synchronization | Reconsider whether the content can be fetched directly; if browser parallelism is genuinely needed, evaluate Selenium Grid or a managed cloud Grid. |
Or skip the browser setup
If the task is to capture a webpage as an image or PDF rather than extract structured text or records, ScreenshotNeo is a screenshot API and MCP server for developers. A single GET request can return PNG, JPEG, WebP, or PDF. It is not a Selenium replacement for scraping page data: use Selenium when you need to inspect and extract DOM content or drive a multi-step browser workflow.
For a screenshot, the cURL call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters and options. Its cookie/consent handling accepts the banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. An MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Yearly billing gives two months free, and every feature is on every plan. Sign up for ScreenshotNeo free to get 1,000 screenshots a month with no card.
Frequently Asked Questions
Can Selenium return the rendered HTML after JavaScript runs?
Yes. Once the page has reached the state your scraper needs, WebDriver can read the current DOM, for example through the driver’s page source. The important part is synchronizing before reading; navigation completion alone may be too early.
Does Selenium bypass a site’s access controls or make scraping automatically permitted?
No. Browser automation does not grant permission or settle whether collection is lawful. Check the target’s rules and applicable requirements before collecting data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

