Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Automation

How to Scrape Dynamic Content with Selenium and Beautiful Soup

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Selenium to render the page, wait for the data you actually need, then pass Selenium’s captured markup to Beautiful Soup. Selenium controls a real browser and executes JavaScript; Beautiful Soup parses the resulting HTML or XML. A reliable scraper therefore follows this sequence: open the URL, wait for a meaningful content condition, read driver.page_source, parse it with an explicitly selected Beautiful Soup parser, and validate the extracted fields.

What Selenium and Beautiful Soup each do

Selenium WebDriver drives a browser: it navigates, executes page JavaScript, clicks controls, fills forms and exposes the browser’s current DOM. Beautiful Soup does not run JavaScript or operate a browser. It turns markup you provide into a searchable parse tree.

That division matters for single-page applications and other sites that insert records after the initial response. A browser can report that navigation is complete while JavaScript is still fetching data or replacing placeholders. Parsing the first response in Beautiful Soup will then find no records, even though a person can see them a moment later.

Install the Python dependencies

Create an isolated environment and install Selenium and Beautiful Soup:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install --upgrade pip
pip install selenium beautifulsoup4

Selenium also needs a compatible browser and driver. Recent Selenium releases can manage drivers for common browsers, but the browser itself must still be installed. Pin and test versions in a production project because browser and library APIs change.

A complete dynamic-content scraper

The following example waits for visible result content, captures the rendered page, and extracts each result. Replace the URL and selectors with the target site’s current structure.

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException

URL = "https://example.com/page"
RESULTS_SELECTOR = ".results"
ITEM_SELECTOR = ".result"

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1440,1200")

try:
    with webdriver.Chrome(options=options) as driver:
        driver.get(URL)
        wait = WebDriverWait(driver, 10)
        wait.until(
            EC.visibility_of_element_located(
                (By.CSS_SELECTOR, RESULTS_SELECTOR)
            )
        )

        # This is the DOM after browser-side JavaScript has run.
        rendered_html = driver.page_source

    soup = BeautifulSoup(rendered_html, "html.parser")
    records = []
    for item in soup.select(ITEM_SELECTOR):
        records.append({
            "text": item.get_text(" ", strip=True),
            "link": (
                item.select_one("a")["href"]
                if item.select_one("a") and item.select_one("a").has_attr("href")
                else None
            ),
        })

    for record in records:
        print(record)
except TimeoutException:
    raise SystemExit(
        f"Timed out waiting for {RESULTS_SELECTOR}; inspect the selector and page state."
    )

visibility_of_element_located is only an example. Choose a condition that proves your data is ready, not merely that navigation finished.

Choose the right wait

Presence, visibility and text

  • Presence: use presence_of_element_located when the node only needs to exist in the DOM.
  • Visibility: use visibility_of_element_located when hidden placeholders should not count.
  • Expected text: use text_to_be_present_in_element when a status label, count or known value signals completion.
  • Custom predicate: wait until a result list has a minimum number of children, a loading indicator disappears, or a particular attribute changes.
def at_least_five_results(driver):
    return len(driver.find_elements(By.CSS_SELECTOR, ".result")) >= 5

WebDriverWait(driver, 15).until(at_least_five_results)

A fixed time.sleep() is a guess. It can be too short on a slow run and waste time on a fast one. A targeted explicit wait polls until the required state or its timeout.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not mix wait strategies casually

Selenium warns that combining implicit and explicit waits can create unpredictable timing. Use one clear strategy; for this workflow, targeted explicit waits are usually easiest to reason about. Keep the timeout finite and report which condition failed.

Pass the rendered markup to Beautiful Soup

After the wait, driver.page_source supplies the browser’s current page markup. Parse it once, then use CSS selectors or normal Beautiful Soup searches:

soup = BeautifulSoup(driver.page_source, "html.parser")
for card in soup.select("article.card"):
    title = card.select_one("h2")
    price = card.select_one(".price")
    print({
        "title": title.get_text(" ", strip=True) if title else None,
        "price": price.get_text(" ", strip=True) if price else None,
    })

Use defensive checks for optional nodes. Do not assume a browser screenshot and the parsed source are identical: responsive layouts, shadow DOM, iframes and client-side state can affect what is visible or where the data lives. If the target is inside an iframe, switch to that frame before reading it. If content is in a shadow root, Selenium’s element APIs may be more appropriate than Beautiful Soup’s document tree.

Select a parser deliberately

Beautiful Soup supports Python’s built-in html.parser, lxml and html5lib. They can construct different trees from malformed markup. Install the parser you choose and name it explicitly so development and production behave consistently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# pip install lxml
soup = BeautifulSoup(rendered_html, "lxml")

Use html.parser for a dependency-light baseline, or select another parser when its handling of the target markup is required. Validate the selected nodes against a saved capture whenever the site changes.

Know when Selenium is unnecessary

If the required data is already in the initial HTML response, a browser adds startup cost and complexity. In that case, an ordinary HTTP client plus Beautiful Soup may be simpler. Conversely, Selenium is justified when JavaScript creates the content, interaction is required, authentication depends on a browser session, or the site needs a rendered state before the data exists. Confirm the site’s terms and access rules before collecting anything.

Make extraction less brittle

Prefer stable signals

  • Target semantic elements, stable IDs, data attributes or documented classes rather than deeply nested positional selectors.
  • Wait for a state that represents the data, such as a non-empty list or a known status, instead of waiting an arbitrary number of seconds.
  • Keep navigation, waiting, parsing and field normalization in separate functions so a selector change has one repair point.
  • Record the URL, timestamp, selector version and number of extracted records for diagnosis.

Handle pagination and “load more” controls

Wait after every click, then parse the new DOM. Stop when the control is absent, disabled or produces no new record identifiers. Deduplicate by a stable URL or site identifier; otherwise repeated renders can create duplicate rows.

Deal with scrolling and lazy content

Some pages request images or records only after an element enters the viewport. Scroll with Selenium, then wait for the resulting records or attribute changes. Do not treat a completed scroll command as proof that network work has finished.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Save evidence on failure

When a wait times out, save driver.page_source and, if useful, a screenshot before closing the driver. Inspect whether the selector changed, a consent dialog blocked the page, authentication failed, or the page displayed a bot challenge. Never silently return an empty dataset as if it were success.

Common failures and fixes

Symptom Likely cause Fix
Beautiful Soup returns no items Parsing occurred before JavaScript inserted them, or the selector is wrong. Wait for the target condition, inspect saved page_source, and verify the selector against the captured DOM.
TimeoutException The condition never became true, the page is slow, or a modal/login blocked it. Check the URL and selector, capture a screenshot and source at timeout, then choose a condition tied to the actual ready state. Increase the timeout only when the state is valid but slow.
Works locally, fails in headless mode Different viewport, timing, browser capabilities or bot defenses. Set an explicit window size, use the same browser version in CI, add diagnostics, and follow the site’s access requirements.
StaleElementReferenceException JavaScript replaced the node after Selenium located it. Wait for the replacement state and locate the element again; avoid holding references across rerenders.
Only part of the list is extracted Infinite scroll or pagination has not completed. Trigger the next batch, wait for a measurable increase, deduplicate, and stop on an explicit end condition.
Malformed or inconsistent fields Parser differences, optional markup or responsive variants. Choose a parser explicitly, use null-safe extraction, normalize whitespace, and test representative captures.

Performance, reliability and cost decisions

Browser automation is heavier than parsing a response because it starts a browser and executes page code. Reuse one driver for a controlled batch, avoid unnecessary waits, and limit images or other resources only when doing so does not remove the data you need. Keep concurrency conservative: many simultaneous browsers can exhaust CPU, memory or the target site’s capacity.

Set page-load and explicit-wait timeouts, retry transient navigation failures with a limit, and make extraction idempotent so a retry does not duplicate records. Cache results where permitted. Treat a timeout, empty result and bot challenge as different outcomes; each requires a different response.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Respect access rules

Check the target’s robots.txt, terms and applicable law before scraping. The Robots Exclusion Protocol describes crawler instructions, but robots rules are not a blanket permission grant or a replacement for legal and policy review. Use a reasonable request rate and avoid disruptive traffic. Handle personal or restricted data according to your obligations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your goal is a clean image or PDF of a rendered page rather than DOM-level field extraction, ScreenshotNeo handles the browser capture through one request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server for AI agents, with take_screenshot, get_page_info and capture_pdf tools.

Here is the one-call cURL form (see the ScreenshotNeo API documentation for options):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python and Node.js clients use the same endpoint:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every plan includes the features: full-page and element capture, device presets or custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting and an OpenAPI specification. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can Beautiful Soup execute JavaScript?

No. It parses markup supplied to it; use Selenium or another browser/runtime to execute page JavaScript first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I parse page_source or an element’s outerHTML?

Use page_source for the rendered document, or capture an element’s outerHTML when you intentionally want one component. In both cases, validate that the captured markup contains the data.

What should I do if the site exposes an API?

Prefer an authorized, documented API when it provides the data you need. It is usually simpler and less resource-intensive than browser automation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.