Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Selenium’s driver.page_source after navigating to the page and waiting for the state you need. In headless Chrome or Firefox, the API is the same as in a visible browser. If you specifically need the browser’s current, JavaScript-mutated DOM, execute document.documentElement.outerHTML instead. Neither method should be treated as a byte-for-byte copy of the original HTTP response.

Choose the kind of HTML you need

“Page source” can mean two different outputs. Selenium’s page_source property returns the WebDriver page-source result for the active browsing context. It is the simplest choice when you want Selenium’s representation of the current page source. A JavaScript call to document.documentElement.outerHTML asks the browser to serialize the live document element after client-side code has changed it.

Goal Use What to expect
Selenium’s page-source result driver.page_source String returned by WebDriver’s page-source command; convenient and browser-independent.
Current DOM after JavaScript mutations driver.execute_script("return document.documentElement.outerHTML;") Serialization of the live document element at the moment the script runs.
Original wire response A network-capture approach Do not assume either Selenium method is the untouched HTTP response body.

Both calls operate in the active window and frame. If the markup is inside an iframe, switch into that frame before capturing it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Selenium and a headless browser

Python package

Create an isolated environment and install Selenium:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
python -m pip install -U selenium

Use a locally installed Chrome or Firefox. Recent Selenium releases can commonly obtain a compatible driver through Selenium Manager; if your environment blocks that behavior, install and expose the matching browser driver yourself. Keep the browser and driver versions compatible.

Headless flags

Chrome and Firefox both support headless operation. The examples below use Chrome’s --headless flag, which works across current Selenium Python setups. On a server without a display, add --no-sandbox and --disable-dev-shm-usage only when your container or host requires them; those flags change browser security and resource behavior and are not universally necessary.

Save page source with Selenium in headless Chrome

This complete script waits for the document to report complete, writes the returned string as UTF-8, and always quits the browser:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait

URL = "https://example.com"

options = Options()
options.add_argument("--headless")
# Add these only if your execution environment needs them:
# options.add_argument("--no-sandbox")
# options.add_argument("--disable-dev-shm-usage")

driver = webdriver.Chrome(options=options)
try:
    driver.get(URL)
    WebDriverWait(driver, 10).until(
        lambda d: d.execute_script("return document.readyState") == "complete"
    )
    html = driver.page_source
    with open("page.html", "w", encoding="utf-8") as output:
        output.write(html)
finally:
    driver.quit()

The official Selenium Python API describes page_source as “Gets the source of the current page.” The property invokes WebDriver’s GET_PAGE_SOURCE command. The call returns a Python string, so writing with UTF-8 preserves Unicode characters in the saved file.

Run it

python save_source.py
# inspect the result
python -c "from pathlib import Path; print(Path('page.html').stat().st_size)"

Open page.html in an editor or browser. Relative links, images, stylesheets, and scripts may not work from a local file because the saved HTML does not copy those dependent resources.

Capture the live DOM after JavaScript runs

Single-page applications often replace or append nodes after the initial response. In that case, serialize the document in the browser:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait

options = Options()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")
    wait = WebDriverWait(driver, 15)
    wait.until(lambda d: d.execute_script("return document.readyState") == "complete")

    live_html = driver.execute_script(
        "return document.documentElement.outerHTML;"
    )
    with open("live-dom.html", "w", encoding="utf-8") as output:
        output.write(live_html)
finally:
    driver.quit()

execute_script runs synchronous JavaScript in the current window. The expression returns the document element’s serialized outerHTML. It is useful when you need mutations already applied to the DOM, but it is still a serialization, not a record of every browser-internal state or network response.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for the content you actually need

document.readyState == "complete" indicates that the document’s standard load lifecycle has completed. It does not prove that an application has finished fetching data. Prefer an explicit, site-specific condition:

from selenium.webdriver.common.by import By

wait.until(lambda d: d.find_element(By.CSS_SELECTOR, "main.results"))
# or wait for a loading marker to disappear
wait.until(lambda d: not d.find_elements(By.CSS_SELECTOR, ".loading"))

Use a selector, text condition, URL change, or another observable readiness signal that represents the state you intend to save. A fixed time.sleep can be useful for a known animation, but it is slower when the page is fast and unreliable when the page is slow. Set a timeout that matches the application’s normal worst case and handle timeout failures explicitly.

Wait for a specific text value

wait.until(
    lambda d: "42 results" in d.find_element(By.CSS_SELECTOR, "main").text
)

After the wait, call either page_source or execute_script. Capture immediately after the condition so later navigation or refresh cannot change the result.

Frames, windows, and shadow DOM

Iframe content

Selenium commands target the active browsing context. Switch to the frame containing the markup before capturing:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from selenium.webdriver.common.by import By

frame = wait.until(lambda d: d.find_element(By.CSS_SELECTOR, "iframe.payment"))
driver.switch_to.frame(frame)
try:
    frame_html = driver.page_source
finally:
    driver.switch_to.default_content()

If you need the outer page and the iframe, capture them separately. The outer document’s HTML does not inline the iframe’s independent document.

Multiple tabs or windows

After opening a new tab, switch to its window handle before calling the source property:

driver.switch_to.window(driver.window_handles[-1])
html = driver.page_source

Shadow DOM

Regular document serialization may not include the internal markup of a shadow root in the way you expect. Query the shadow root with JavaScript and serialize that specific node when the component exposes one:

shadow_html = driver.execute_script("""
const host = document.querySelector('my-component');
return host && host.shadowRoot ? host.shadowRoot.innerHTML : null;
""")

Closed shadow roots cannot be inspected through ordinary page JavaScript.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make the capture reproducible

  • Set a deterministic viewport with options.add_argument("--window-size=1440,900") when responsive breakpoints affect the markup.
  • Use an explicit user agent only when your test requires one; changing it can select different server responses.
  • Record the final URL, timestamp, browser version, and selected frame alongside the HTML.
  • Write to a temporary file and rename it after a successful write when another process consumes the output.
  • Always call driver.quit() in a finally block so headless browser processes do not accumulate.

Common failures and fixes

“Unable to obtain driver” or session creation failure

Check that Chrome or Firefox is installed and that the driver is compatible. Upgrade Selenium, allow Selenium Manager to reach its driver source, or install a matching driver and put it on PATH. In a container, verify executable permissions and required shared libraries.

The saved HTML is missing data

The capture probably happened before the asynchronous request finished. Replace the document-ready wait with a wait for the result selector, expected text, or disappearance of the loading element. Increase the timeout only after choosing a meaningful condition.

The output contains an old version of the page

Confirm that navigation completed, that you are on the intended window and frame, and that the application did not navigate again after your wait. Log driver.current_url immediately before capture.

An iframe’s markup is absent

Switch to the iframe first. Cross-origin frames remain separate browsing contexts; capture their content from within the frame if the page and browser permissions allow it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Headless mode behaves differently

Set a realistic window size, check responsive breakpoints, and compare browser and driver versions. Some sites deliver bot challenges or different content to automated browsers; Selenium source capture cannot bypass a challenge reliably.

Timeout while waiting

Inspect whether the selector exists in this viewport and whether a consent dialog or login gate prevents the expected state. Capture diagnostic data such as the current URL and a screenshot before quitting, then fix the prerequisite rather than replacing the wait with an arbitrary long sleep.

HTML encoding looks corrupted

Write the Python string with encoding="utf-8". If another system reads the file, ensure it also treats the file as UTF-8 and that your editor is not guessing a legacy encoding.

Performance, reliability, and limits

Starting a browser is substantially heavier than requesting a static URL, so reuse one driver for a batch of pages when isolation requirements permit. Navigate sequentially, wait only for the state required by each page, and quit the driver at the end of the batch. A page with large images, long-running scripts, or infinite scrolling can consume considerable memory even though the final HTML is text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For reliable jobs, give each navigation a timeout, catch WebDriver exceptions, and save a failure record containing the URL and error. Do not treat an empty or unusually short document as success without checking the page’s expected selector. If you need the server’s exact response bytes, collect network traffic with a method designed for that requirement; Selenium’s page-source result and live-DOM serialization answer different questions.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo provides a website-capture API when your deliverable is a rendered screenshot or PDF rather than HTML source. A single GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options and response details. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Every plan includes the features. The Free plan provides 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Other options include full-page captures with lazy images loaded, CSS-selector element capture, device presets and custom viewports, dark mode, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for 1,000 free screenshots a month with no card.

Frequently asked questions

Does headless mode change the Selenium API?

No. Add the browser’s headless option, then read driver.page_source or execute the same JavaScript as you would in a visible session.

Can page source include JavaScript-generated content?

It can reflect content present when WebDriver returns the page-source result, but for an explicit live-DOM serialization use document.documentElement.outerHTML after waiting for the application’s readiness condition.

How do I get only one element’s HTML?

Locate the element and read its get_attribute("outerHTML"), or execute JavaScript that returns that element’s outerHTML. This avoids saving the entire document when a component is all you need.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is the returned source guaranteed to match “View Source”?

No. “View Source” represents the browser’s source view, while WebDriver’s page-source command and live-DOM serialization have their own semantics. Choose the output based on whether you need a WebDriver result, current DOM, or raw network response.

Frequently Asked Questions

Can I capture page source without displaying a browser window?

Yes. Configure Chrome or Firefox with its headless option; the retrieval calls remain unchanged.

Should I use a fixed sleep before reading page_source?

Prefer an explicit wait for the selector or state that proves the content you need is ready. Fixed sleeps are timing guesses.

What happens if the page requires login?

Authenticate in the same WebDriver session, then wait for a post-login selector before capturing. Do not put credentials in source files or command-line logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.