What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use Selenium to render the page, wait for the data you actually need, then pass Selenium’s captured markup to Beautiful Soup. Selenium controls a real browser and executes JavaScript; Beautiful Soup parses the resulting HTML or XML. A reliable scraper therefore follows this sequence: open the URL, wait for a meaningful content condition, read driver.page_source, parse it with an explicitly selected Beautiful Soup parser, and validate the extracted fields.
What Selenium and Beautiful Soup each do
Selenium WebDriver drives a browser: it navigates, executes page JavaScript, clicks controls, fills forms and exposes the browser’s current DOM. Beautiful Soup does not run JavaScript or operate a browser. It turns markup you provide into a searchable parse tree.
That division matters for single-page applications and other sites that insert records after the initial response. A browser can report that navigation is complete while JavaScript is still fetching data or replacing placeholders. Parsing the first response in Beautiful Soup will then find no records, even though a person can see them a moment later.
Install the Python dependencies
Create an isolated environment and install Selenium and Beautiful Soup:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install --upgrade pip
pip install selenium beautifulsoup4
Selenium also needs a compatible browser and driver. Recent Selenium releases can manage drivers for common browsers, but the browser itself must still be installed. Pin and test versions in a production project because browser and library APIs change.
A complete dynamic-content scraper
The following example waits for visible result content, captures the rendered page, and extracts each result. Replace the URL and selectors with the target site’s current structure.
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException
URL = "https://example.com/page"
RESULTS_SELECTOR = ".results"
ITEM_SELECTOR = ".result"
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1440,1200")
try:
with webdriver.Chrome(options=options) as driver:
driver.get(URL)
wait = WebDriverWait(driver, 10)
wait.until(
EC.visibility_of_element_located(
(By.CSS_SELECTOR, RESULTS_SELECTOR)
)
)
# This is the DOM after browser-side JavaScript has run.
rendered_html = driver.page_source
soup = BeautifulSoup(rendered_html, "html.parser")
records = []
for item in soup.select(ITEM_SELECTOR):
records.append({
"text": item.get_text(" ", strip=True),
"link": (
item.select_one("a")["href"]
if item.select_one("a") and item.select_one("a").has_attr("href")
else None
),
})
for record in records:
print(record)
except TimeoutException:
raise SystemExit(
f"Timed out waiting for {RESULTS_SELECTOR}; inspect the selector and page state."
)
visibility_of_element_located is only an example. Choose a condition that proves your data is ready, not merely that navigation finished.
Choose the right wait
Presence, visibility and text
- Presence: use
presence_of_element_locatedwhen the node only needs to exist in the DOM. - Visibility: use
visibility_of_element_locatedwhen hidden placeholders should not count. - Expected text: use
text_to_be_present_in_elementwhen a status label, count or known value signals completion. - Custom predicate: wait until a result list has a minimum number of children, a loading indicator disappears, or a particular attribute changes.
def at_least_five_results(driver):
return len(driver.find_elements(By.CSS_SELECTOR, ".result")) >= 5
WebDriverWait(driver, 15).until(at_least_five_results)
A fixed time.sleep() is a guess. It can be too short on a slow run and waste time on a fast one. A targeted explicit wait polls until the required state or its timeout.
Do not mix wait strategies casually
Selenium warns that combining implicit and explicit waits can create unpredictable timing. Use one clear strategy; for this workflow, targeted explicit waits are usually easiest to reason about. Keep the timeout finite and report which condition failed.
Pass the rendered markup to Beautiful Soup
After the wait, driver.page_source supplies the browser’s current page markup. Parse it once, then use CSS selectors or normal Beautiful Soup searches:
soup = BeautifulSoup(driver.page_source, "html.parser")
for card in soup.select("article.card"):
title = card.select_one("h2")
price = card.select_one(".price")
print({
"title": title.get_text(" ", strip=True) if title else None,
"price": price.get_text(" ", strip=True) if price else None,
})
Use defensive checks for optional nodes. Do not assume a browser screenshot and the parsed source are identical: responsive layouts, shadow DOM, iframes and client-side state can affect what is visible or where the data lives. If the target is inside an iframe, switch to that frame before reading it. If content is in a shadow root, Selenium’s element APIs may be more appropriate than Beautiful Soup’s document tree.
Select a parser deliberately
Beautiful Soup supports Python’s built-in html.parser, lxml and html5lib. They can construct different trees from malformed markup. Install the parser you choose and name it explicitly so development and production behave consistently.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
# pip install lxml
soup = BeautifulSoup(rendered_html, "lxml")
Use html.parser for a dependency-light baseline, or select another parser when its handling of the target markup is required. Validate the selected nodes against a saved capture whenever the site changes.
Know when Selenium is unnecessary
If the required data is already in the initial HTML response, a browser adds startup cost and complexity. In that case, an ordinary HTTP client plus Beautiful Soup may be simpler. Conversely, Selenium is justified when JavaScript creates the content, interaction is required, authentication depends on a browser session, or the site needs a rendered state before the data exists. Confirm the site’s terms and access rules before collecting anything.
Make extraction less brittle
Prefer stable signals
- Target semantic elements, stable IDs, data attributes or documented classes rather than deeply nested positional selectors.
- Wait for a state that represents the data, such as a non-empty list or a known status, instead of waiting an arbitrary number of seconds.
- Keep navigation, waiting, parsing and field normalization in separate functions so a selector change has one repair point.
- Record the URL, timestamp, selector version and number of extracted records for diagnosis.
Handle pagination and “load more” controls
Wait after every click, then parse the new DOM. Stop when the control is absent, disabled or produces no new record identifiers. Deduplicate by a stable URL or site identifier; otherwise repeated renders can create duplicate rows.
Deal with scrolling and lazy content
Some pages request images or records only after an element enters the viewport. Scroll with Selenium, then wait for the resulting records or attribute changes. Do not treat a completed scroll command as proof that network work has finished.
Save evidence on failure
When a wait times out, save driver.page_source and, if useful, a screenshot before closing the driver. Inspect whether the selector changed, a consent dialog blocked the page, authentication failed, or the page displayed a bot challenge. Never silently return an empty dataset as if it were success.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Beautiful Soup returns no items | Parsing occurred before JavaScript inserted them, or the selector is wrong. | Wait for the target condition, inspect saved page_source, and verify the selector against the captured DOM. |
| TimeoutException | The condition never became true, the page is slow, or a modal/login blocked it. | Check the URL and selector, capture a screenshot and source at timeout, then choose a condition tied to the actual ready state. Increase the timeout only when the state is valid but slow. |
| Works locally, fails in headless mode | Different viewport, timing, browser capabilities or bot defenses. | Set an explicit window size, use the same browser version in CI, add diagnostics, and follow the site’s access requirements. |
| StaleElementReferenceException | JavaScript replaced the node after Selenium located it. | Wait for the replacement state and locate the element again; avoid holding references across rerenders. |
| Only part of the list is extracted | Infinite scroll or pagination has not completed. | Trigger the next batch, wait for a measurable increase, deduplicate, and stop on an explicit end condition. |
| Malformed or inconsistent fields | Parser differences, optional markup or responsive variants. | Choose a parser explicitly, use null-safe extraction, normalize whitespace, and test representative captures. |
Performance, reliability and cost decisions
Browser automation is heavier than parsing a response because it starts a browser and executes page code. Reuse one driver for a controlled batch, avoid unnecessary waits, and limit images or other resources only when doing so does not remove the data you need. Keep concurrency conservative: many simultaneous browsers can exhaust CPU, memory or the target site’s capacity.
Set page-load and explicit-wait timeouts, retry transient navigation failures with a limit, and make extraction idempotent so a retry does not duplicate records. Cache results where permitted. Treat a timeout, empty result and bot challenge as different outcomes; each requires a different response.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Respect access rules
Check the target’s robots.txt, terms and applicable law before scraping. The Robots Exclusion Protocol describes crawler instructions, but robots rules are not a blanket permission grant or a replacement for legal and policy review. Use a reasonable request rate and avoid disruptive traffic. Handle personal or restricted data according to your obligations.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Best Value
Or skip the browser setup
If your goal is a clean image or PDF of a rendered page rather than DOM-level field extraction, ScreenshotNeo handles the browser capture through one request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server for AI agents, with take_screenshot, get_page_info and capture_pdf tools.
Here is the one-call cURL form (see the ScreenshotNeo API documentation for options):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python and Node.js clients use the same endpoint:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every plan includes the features: full-page and element capture, device presets or custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting and an OpenAPI specification. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently Asked Questions
Can Beautiful Soup execute JavaScript?
No. It parses markup supplied to it; use Selenium or another browser/runtime to execute page JavaScript first.
Should I parse page_source or an element’s outerHTML?
Use page_source for the rendered document, or capture an element’s outerHTML when you intentionally want one component. In both cases, validate that the captured markup contains the data.
What should I do if the site exposes an API?
Prefer an authorized, documented API when it provides the data you need. It is usually simpler and less resource-intensive than browser automation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




