Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Selenium when the data appears only after a browser runs JavaScript or when you must reproduce clicks, scrolling, login, or other user actions. A reliable Python scraper creates a WebDriver session, opens the page, waits for a specific condition, locates elements with stable selectors, extracts text or attributes, and always closes the browser. The complete example below handles a JavaScript-rendered product list and includes timeout, selector, and cleanup practices you can adapt to another site.
What Selenium adds to screen scraping
Selenium WebDriver drives a browser natively. Unlike a direct HTTP client, it can execute the page’s JavaScript and expose the rendered state that was not present in the initial HTML. That makes it useful for single-page applications, infinite-scroll pages, search interfaces, and workflows that require interaction.
A successful driver.get() call means the browser reached the page’s load event according to its page-load strategy. It does not prove that an AJAX request has finished or that a JavaScript-created element is ready. Synchronization is therefore the central design problem in Selenium scraping.
Free tools Windows power users keep installed
One-click scans. No signup required.
Before collecting anything, check the target’s published API, terms, authentication requirements, robots guidance, and rate limits. Prefer an official API or a lightweight HTTP request when it provides the data you need; a browser is slower and consumes substantially more memory and CPU than a direct request.
#1 Best Overall
Install Python and Selenium
- Use a supported Python installation and create an isolated virtual environment:
python -m venv .venv # macOS/Linux source .venv/bin/activate # Windows PowerShell .venvScriptsActivate.ps1 - Install or upgrade the Python binding:
python -m pip install --upgrade selenium - Save the script as
scrape.pyand run it withpython scrape.py.
Recent Selenium releases can normally obtain a compatible browser driver automatically through Selenium Manager. If your environment blocks that process, install a driver compatible with the browser version and make it available on PATH, or pass an explicit service object.
A complete scraper for a JavaScript-rendered page
This example waits for product cards, extracts their title, price, and link, and writes JSON. Replace the URL and CSS selectors with selectors from the site you are allowed to collect.
import json
from selenium import webdriver
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URL = "https://example.com/products"
WAIT_SECONDS = 15
options = webdriver.ChromeOptions()
# Keep the browser visible while developing. Add this for a server:
# options.add_argument("--headless=new")
options.add_argument("--window-size=1440,1000")
# Selenium Manager usually finds the matching driver.
driver = webdriver.Chrome(options=options)
driver.set_page_load_timeout(30)
driver.set_script_timeout(30)
# The default implicit timeout is zero. Keep it explicit rather than mixing
# implicit and explicit waits.
driver.implicitly_wait(0)
try:
driver.get(URL)
wait = WebDriverWait(driver, WAIT_SECONDS, poll_frequency=0.5)
# Wait for the rendered cards, not merely document.readyState.
cards = wait.until(
EC.visibility_of_all_elements_located(
(By.CSS_SELECTOR, "article.product-card")
)
)
records = []
for card in cards:
title = card.find_element(By.CSS_SELECTOR, ".product-title").text.strip()
price = card.find_element(By.CSS_SELECTOR, ".price").text.strip()
link = card.find_element(By.CSS_SELECTOR, "a").get_attribute("href")
records.append({"title": title, "price": price, "url": link})
with open("products.json", "w", encoding="utf-8") as output:
json.dump(records, output, ensure_ascii=False, indent=2)
print(f"Saved {len(records)} records")
except TimeoutException:
print("Timed out waiting for product cards")
print("Current URL:", driver.current_url)
print("Page title:", driver.title)
raise
finally:
driver.quit()
The finally block matters: it closes the browser even when navigation, extraction, or a wait raises an exception. Leaving sessions open eventually exhausts resources in a long-running job.
Choose locators that survive redesigns
The Python bindings support ID, name, XPath, link text, partial link text, tag name, class name, and CSS selector strategies. Prefer the most stable attribute the site exposes and scope the search to the smallest useful container.
Recommended order
- Unique ID:
By.ID, "results"when the ID is stable and not generated per session. - Semantic or test attribute: for example,
By.CSS_SELECTOR, '[data-testid="product-card"]'when the site documents or consistently uses it. - Stable CSS structure: a component class scoped under its parent, such as
article.product-card .price. - XPath: useful for relationships or text conditions, but avoid long absolute paths such as
/html/body/div[2]/div[1]. - Link text: convenient for a known link, but fragile when wording or localization changes.
find_element returns the first match and raises an exception if none exists. find_elements returns a list, including an empty list when there are no matches. Use the latter when zero results is a valid outcome, and scope nested searches to each card so a page header does not accidentally supply a value.
Rank #2
Wait for the state you actually need
Explicit waits poll a condition until it succeeds or the timeout expires. The documented default polling interval for WebDriverWait is 0.5 seconds.
Common conditions
from selenium.webdriver.support import expected_conditions as EC
# Exists in the DOM (may still be hidden)
wait.until(EC.presence_of_element_located((By.ID, "results")))
# Visible to the user
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, ".result")))
# Ready for a click
wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.next")))
# A specific condition you define
wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, ".result")) >= 20)
Wait for a meaningful application state: a result count, a loading indicator disappearing, a button becoming enabled, or a required card appearing. Arbitrary time.sleep() calls either waste time on fast runs or fail on slow ones.
Do not mix implicit and explicit waits. Selenium warns that combining them can produce unpredictable, longer delays because each explicit poll may include the implicit timeout. Set the implicit timeout deliberately—zero is a clear choice when all synchronization is explicit.
Interactions before extraction
Clicking a “load more” control
while True:
old_count = len(driver.find_elements(By.CSS_SELECTOR, "article.product-card"))
try:
button = WebDriverWait(driver, 5).until(
EC.element_to_be_clickable((By.CSS_SELECTOR, "button.load-more"))
)
except TimeoutException:
break
driver.execute_script("arguments[0].click();", button)
WebDriverWait(driver, 10).until(
lambda d: len(d.find_elements(By.CSS_SELECTOR, "article.product-card")) > old_count
)
Stop when the control disappears, becomes disabled, or the item count stops increasing. Add a maximum-page or maximum-item limit so a faulty endpoint cannot create an endless loop.
Scrolling an infinite list
previous = 0
for _ in range(20):
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
WebDriverWait(driver, 10).until(
lambda d: len(d.find_elements(By.CSS_SELECTOR, "article.product-card")) > previous
)
previous = len(driver.find_elements(By.CSS_SELECTOR, "article.product-card"))
Some applications virtualize rows, removing off-screen nodes. In that case, extract each batch before scrolling onward rather than assuming the entire list remains in the DOM.
Reading attributes and page source
Use element.text for visible text and get_attribute("href"), get_attribute("content"), or another attribute when the value is not rendered as text. driver.page_source gives the current serialized DOM, which is useful for diagnostics but is not a substitute for waiting for the correct state.
Timeouts, page-load strategy, and browser configuration
Configure navigation and script limits separately from element waits:
driver.set_page_load_timeout(30)
driver.set_script_timeout(30)
driver.implicitly_wait(0)
A page-load strategy can return control earlier, but it also shifts responsibility for readiness to your explicit conditions. Do not treat an early return as proof that the application is usable. Keep browser options close to the workload: headless mode for unattended servers, a realistic window size for responsive layouts, and a fixed user-agent only when you have a legitimate compatibility reason.
When Selenium is the wrong tool
| Need | Better first choice | Why |
|---|---|---|
| Structured data exposed by a documented endpoint | Official API | Usually more stable, cheaper, and authorized for automation. |
| Static HTML with no interaction | HTTP client plus an HTML parser | No browser startup or synchronization overhead. |
| JavaScript rendering, clicks, login, or scrolling are essential | Selenium | It reproduces browser behavior and user flows. |
Also compare authentication requirements, rate limits, locator maintenance, observability, and debugging effort. Browser automation is not a way around access controls; bot checks and CAPTCHAs indicate that you should stop and review the site’s rules rather than attempt to defeat them.
Troubleshooting common failures
“Unable to obtain driver” or browser/driver mismatch
Upgrade Selenium and the browser, allow Selenium Manager to reach its driver sources, or install a matching driver explicitly. Check that the executable is on PATH and that your process has permission to launch it.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
TimeoutException while the page looks complete
Inspect the selector in browser developer tools, verify the element is inside an iframe, and check whether a consent dialog or login step blocks rendering. If the content is inside an iframe, switch first:
frame = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "iframe")))
driver.switch_to.frame(frame)
# locate elements inside the frame
driver.switch_to.default_content()
Replace a fixed delay with a condition tied to the site’s actual state. Capture driver.current_url, driver.title, and a screenshot or page source on failure.
Element is present but not clickable
Presence only proves that a node exists. Wait for visibility or clickability, scroll it into view, and check for an overlay, disabled state, or stale reference. Re-find elements after the page re-renders instead of reusing an old reference.
Empty text or missing attributes
The value may be added later, stored in a property rather than an attribute, or supplied by a pseudo-element. Wait for a non-empty value, inspect the rendered DOM, and use the site’s underlying API only when its terms and authentication permit it.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteMemory growth in batch jobs
Reuse one session only when the site’s state can safely persist; otherwise quit between jobs. Bound queues, limit concurrency, avoid downloading unnecessary resources where appropriate, and record failures rather than retrying indefinitely.
Best Value
Make runs reproducible and respectful
- Record the URL, timestamp, selector version, browser version, and outcome for every job.
- Use a conservative request rate and exponential backoff for transient navigation failures.
- Deduplicate records by a stable site identifier or canonical URL.
- Validate required fields before writing output; keep raw HTML or screenshots only when your retention policy allows it.
- Handle login credentials through environment variables or a secret manager, never source code.
- Test selectors against representative layouts, including empty results, pagination boundaries, and localized pages.
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than DOM-level data extraction, ScreenshotNeo provides a single screenshot API request. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off.
Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
Here is the one-call cURL version (see the ScreenshotNeo documentation for all options):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', body);
Every plan includes the features: full-page and element capture, device presets and custom viewports, dark mode, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. The parameter names used by other screenshot APIs also work for easier migration.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free. Create a free ScreenshotNeo account to start.
Frequently Asked Questions
Can Selenium scrape content rendered inside a shadow DOM?
Sometimes. Standard selectors cannot cross a closed shadow root; for an open root, access the shadow root with Selenium’s shadow-DOM support or execute JavaScript, then locate elements within that root.
Should I save the browser’s cookies between runs?
Only when the site’s permitted workflow requires a persistent session. Store cookies securely, scope them to the intended domain, and avoid carrying authenticated state into unrelated jobs.
Recommended Free Tools
Is Selenium suitable for high-volume crawling?
It can be distributed, but each browser session is relatively resource-intensive. For high volume, first look for an authorized API or direct HTTP workflow, then add strict concurrency, rate, and retry limits if browser automation remains necessary.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

