Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Yes—Selenium can scrape JavaScript-rendered sites with Python. It launches a real browser, waits for the application state you need, and then reads the same DOM a user sees. This guide builds a small scraper that collects article titles and links, then expands it with robust locators, explicit waits, pagination, sessions, retries, checkpoints, and responsible-use safeguards.
What you will build
Our example collects article cards from a page whose content is rendered or updated by JavaScript. The pattern applies to product listings, dashboards, search results, and other sites where a plain HTTP request returns incomplete HTML.
- Python 3.10 or newer.
- A current installation of Chrome, Edge, Firefox, Safari, WebKitGTK, or WPEWebKit.
- A virtual environment and the Selenium Python package.
Current Selenium Python documentation supports Python 3.10+ and Selenium 4.49.0 documentation describes browser automation across those browser families. Modern Selenium includes Selenium Manager, which usually obtains a compatible driver automatically.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →1. Create an isolated project
mkdir selenium-scraper
cd selenium-scraper
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install -U selenium
Keep your scraper and its output in this directory. A lock file or pinned package version is sensible for production, but update Selenium regularly so browser compatibility fixes are included.
#1 Best Overall
2. Launch a browser and inspect the page
Start with webdriver.Chrome(). Selenium Manager handles driver discovery in common cases, so a separate ChromeDriver download is not normally required.
from selenium import webdriver
browser = webdriver.Chrome()
try:
browser.get("https://example.com")
print(browser.title)
print(browser.current_url)
finally:
browser.quit()
Replace the URL with a page you are permitted to access. Inspect it in your browser’s developer tools: right-click a target, choose Inspect, and identify the smallest stable element that contains the data. Examine the live DOM after JavaScript has run, not only the original page source.
3. Choose locators that survive redesigns
Prefer a stable ID
Selenium’s locator guidance says: “In general, if HTML IDs are available, unique, and consistently predictable, they are the preferred method for locating an element on a page.” An ID such as article-list is usually clearer and less fragile than a long class chain.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteUse compact CSS next
When no reliable ID exists, use a short CSS selector such as article.card h2 a. Avoid presentation-only classes, generated names, and selectors that depend on an exact nesting depth.
Use XPath for relationships
XPath is useful when you must locate an element by relationship or text, but Selenium describes it as harder to debug and typically slower. Keep it narrow and readable, for example //article[.//h2]/descendant::a[1].
Rank #2
| Locator | Best use | Main risk |
|---|---|---|
| Unique ID | Stable, direct lookup | Not present or regenerated |
| CSS selector | Readable combinations and descendants | Breaks when classes are renamed |
| XPath | Text and structural relationships | Harder maintenance and debugging |
4. Wait for the state you actually need
driver.get() returning only tells you that the selected page-load policy has completed; it does not prove that an XHR, fetch call, click, infinite-scroll operation, or client-side route has populated your target.
Use an explicit wait
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
wait = WebDriverWait(browser, 10)
list_element = wait.until(
EC.presence_of_element_located((By.CSS_SELECTOR, "article"))
)
wait.until(EC.visibility_of(list_element))
Presence means the node exists in the DOM; visibility also requires it to be displayed. For dynamic content, add a condition that proves the right data arrived:
wait.until(
EC.text_to_be_present_in_element(
(By.CSS_SELECTOR, "article h2"), "Python"
)
)
Prefer condition-based synchronization to arbitrary time.sleep(). A fixed sleep either under-waits on a slow run or wastes time on a fast one.
Do not combine wait policies accidentally
Selenium’s waiting guidance states exactly: “Do not mix implicit and explicit waits.” An implicit wait changes how every element lookup polls, making explicit timeout behavior difficult to predict. Choose explicit waits for a scraper whose states differ by page or component.
5. A complete scraper
This script waits for cards, extracts text and links, follows a “Next” control until it disappears, retries transient page failures, and checkpoints each page to JSON Lines. Replace the selectors with those from your target DOM.
import json
import time
from pathlib import Path
from urllib.parse import urljoin
from selenium import webdriver
from selenium.common.exceptions import TimeoutException, WebDriverException
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
START_URL = "https://example.com/news"
CARD = "article.card"
NEXT = "a[rel='next']"
OUTPUT = Path("articles.jsonl")
options = webdriver.ChromeOptions()
# options.add_argument("--headless=new") # enable for unattended runs
options.page_load_strategy = "normal" # normal, eager, or none
browser = webdriver.Chrome(options=options)
browser.set_page_load_timeout(45)
browser.set_script_timeout(30)
wait = WebDriverWait(browser, 15)
def read_cards():
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, CARD)))
rows = []
for card in browser.find_elements(By.CSS_SELECTOR, CARD):
heading = card.find_element(By.CSS_SELECTOR, "h2")
link = heading.find_element(By.CSS_SELECTOR, "a")
href = link.get_attribute("href")
rows.append({
"title": heading.text.strip(),
"url": urljoin(browser.current_url, href),
})
return rows
def next_page():
try:
control = browser.find_element(By.CSS_SELECTOR, NEXT)
if not control.is_displayed() or not control.is_enabled():
return False
old_url = browser.current_url
browser.execute_script("arguments[0].click();", control)
wait.until(lambda d: d.current_url != old_url or
len(d.find_elements(By.CSS_SELECTOR, CARD)) > 0)
return True
except (TimeoutException, WebDriverException):
return False
try:
browser.get(START_URL)
seen = set()
with OUTPUT.open("w", encoding="utf-8") as stream:
for page_number in range(1, 101):
for attempt in range(3):
try:
records = read_cards()
break
except (TimeoutException, WebDriverException):
if attempt == 2:
raise
time.sleep(2 ** attempt)
new_records = [r for r in records if r["url"] not in seen]
for record in new_records:
seen.add(record["url"])
record["page"] = page_number
stream.write(json.dumps(record, ensure_ascii=False) + "n")
stream.flush() # checkpoint survives an interrupted run
if not next_page():
break
finally:
browser.quit()
6. Extract more than visible text
Attributes and links
Use get_attribute() for URLs, image sources, data attributes, and ARIA labels. Resolve relative links with urljoin(), as the example does.
Tables
for row in browser.find_elements(By.CSS_SELECTOR, "table tbody tr"):
cells = [cell.text.strip() for cell in
row.find_elements(By.CSS_SELECTOR, "th, td")]
print(cells)
Shadow DOM and iframes
For an iframe, switch first with browser.switch_to.frame(frame_element), scrape its document, then call browser.switch_to.default_content(). Components in an open shadow root require Selenium’s shadow-root APIs; closed roots are not directly queryable from page JavaScript.
7. Page-load strategy, timeouts, and browser options
| Strategy | Browser returns after | Use when |
|---|---|---|
normal |
The load event and dependent resources finish | Safest default |
eager |
DOMContentLoaded | Images and nonessential assets are irrelevant |
none |
Navigation starts without blocking | You have strong, explicit readiness checks |
Set page-load, script, and element timeouts deliberately. A short page-load timeout does not replace an explicit wait for the card that appears after an API call. A proxy can be configured through browser options when a restricted network, traffic capture, or mock backend requires it.
8. Pagination, sessions, and polite throughput
- Preserve the same driver so cookies, login state, and local storage remain available.
- Stop on a missing, disabled, or unchanged “Next” control; also deduplicate URLs to prevent loops.
- Retry transient navigation failures with a small capped backoff, as in the example. Do not retry indefinitely.
- Flush checkpoints after each page so a crash does not discard earlier work.
- For infinite scroll, scroll in measured increments and wait for the item count to increase before continuing.
- Limit concurrency and add delays appropriate to the site. A real browser consumes substantially more CPU, memory, and bandwidth than a static HTTP client.
9. Diagnose common failures
“NoSuchElementException”
The selector may be wrong, the element may be inside an iframe or shadow root, or the app has not rendered it. Inspect the post-JavaScript DOM, switch into the correct frame, and wait for a meaningful condition before lookup.
“Element not interactable” or click interception
A modal, sticky header, animation, or consent layer may cover the control. Wait for clickability, scroll it into view, close the overlay when permitted, and verify that the click changed the expected state. JavaScript clicking should be a last resort because it can bypass normal user-event behavior.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
TimeoutException
Check whether the selector is stale, the request failed, a bot check appeared, or the page uses a different route than expected. Capture a screenshot and page source on failure, then increase a timeout only after identifying the slow state.
Driver or browser version mismatch
Update the browser and Selenium package, allow Selenium Manager to resolve the driver, and remove an obsolete driver executable from your PATH. Manual driver installation is still appropriate in locked-down environments; match the browser’s major version and point the service explicitly.
Blank or partial data
Confirm that your wait checks the content rather than merely document.readyState. Increase the script timeout for long client operations, check network access, and verify that lazy-loaded items were actually scrolled into view.
10. Responsible scraping
Read the site’s terms and access rules before automating. Inspect robots.txt; RFC 9309 is the IETF reference for the Robots Exclusion Protocol. Robots rules are an access signal, not a blanket legal determination, so obtain permission where required and stop when a site blocks automation. Identify your user agent where appropriate, respect rate limits, and avoid collecting personal data you do not need.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsOr skip the browser setup
If you only need a clean image or PDF rather than DOM-level extraction, ScreenshotNeo provides a single-call website screenshot API. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for all options. The same endpoint supports full-page and element captures, device and viewport settings, retina scale, PDFs, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data, and an OpenAPI specification. Parameters commonly used by other screenshot APIs also work, easing migration.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account to try it.
Best Value
FAQ
Do I still need ChromeDriver?
Usually not: Selenium Manager handles common driver installation. Manual drivers remain useful when policy, networking, or a pinned browser image prevents automatic resolution.
Which wait should I choose for a single scraper?
Use explicit waits tied to the state you need and avoid mixing them with implicit waits. This keeps each page’s synchronization visible and predictable.
When is a static HTTP client better?
Use one when the needed data is present in the initial HTML or a documented API. It is lighter and faster; choose Selenium when JavaScript execution, user interactions, or browser session behavior is essential.
Frequently Asked Questions
Can Selenium scrape content loaded only after scrolling?
Yes. Scroll incrementally, then explicitly wait for the item count or a sentinel element to change before extracting the newly loaded nodes.
How should I make a scraper restartable?
Write each completed record or page to a checkpoint file, deduplicate by a stable URL or ID, and resume from the last confirmed page instead of rebuilding the entire run.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

