Choose the artifact before you write code. ChromeDriver can save the live, post-JavaScript DOM; package a page and its dependencies as one MHTML file; collect individual network responses; or wait for a normal browser download. These outputs are not interchangeable. The examples below use Selenium with Chrome in headless mode, explain the limits of each method, and show how to avoid losing dynamic content or unfinished downloads.
Decide what “save the page” means
| What you need | Use | What you get | Boundary |
|---|---|---|---|
| Current rendered markup | document.documentElement.outerHTML or Chrome --dump-dom |
Serialized DOM after scripts have modified it | It is not the original HTTP response and does not embed external images, CSS, fonts or scripts. |
| One-file archive | DevTools Protocol Page.captureSnapshot or the pageCapture extension API |
MHTML containing the document and captured dependencies | Protocol and extension availability depends on the installed Chrome version; verify it in your deployment. |
| Separate resources or response analysis | Network tracking or ChromeDriver performance logs | Request/response events and, while available, response bodies | You must handle redirects, duplicate URLs, encodings, naming and large bodies. |
| A file linked by a download button | Chrome download preferences plus a completion check | The browser’s downloaded file | ChromeDriver does not wait for completion automatically. |
| Visual or printable output | Screenshot or PDF commands | PNG/JPEG/WebP or PDF | Neither is an HTML or resource archive. |
Prepare compatible Chrome and ChromeDriver
ChromeDriver is Chrome’s WebDriver control layer. In Selenium, pass --headless through Chrome options. Keep the browser and driver compatible. For Chrome 115 and later, Chrome for Testing publishes release-channel binaries and availability information; use that source when you provision pinned CI images. Chrome’s unified Headless implementation changed in Chrome 112, and the former separate implementation moved to the chrome-headless-shell binary beginning with Chrome 132.0.6793.0.
Install Selenium in the environment that will run the capture:
python -m pip install -U selenium
The following helper creates a headless session, uses a dedicated download directory, and avoids assuming that navigation completion means an application has finished fetching data.
#1 Best Overall
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
def make_driver(download_dir=None, performance_logs=False):
options = Options()
options.add_argument("--headless")
options.add_argument("--window-size=1440,1200")
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
options.add_argument("--disable-gpu")
if download_dir:
path = str(Path(download_dir).resolve())
options.add_experimental_option("prefs", {
"download.default_directory": path,
"download.prompt_for_download": False,
"download.directory_upgrade": True,
"safebrowsing.enabled": True,
})
if performance_logs:
options.set_capability("goog:loggingPrefs", {
"performance": "ALL",
"browser": "ALL",
})
return webdriver.Chrome(options=options)
Use an absolute, writable path for downloads. In containers, also ensure the user running Chrome can create and rename files there.
Save the rendered HTML (the live DOM)
This is the right choice when you need the markup a user would see after JavaScript has run. It is a serialization of the current DOM, not a byte-for-byte copy of the server response.
from pathlib import Path
from selenium.webdriver.support.ui import WebDriverWait
url = "https://example.com/app"
out = Path("rendered.html")
driver = make_driver()
try:
driver.get(url)
WebDriverWait(driver, 30).until(
lambda d: d.execute_script("return document.readyState") == "complete"
)
# Replace this with a condition specific to your application.
WebDriverWait(driver, 30).until(
lambda d: d.find_element("css selector", "main[data-loaded='true']")
)
markup = driver.execute_script(
"return document.documentElement.outerHTML;"
)
out.write_text(markup, encoding="utf-8")
finally:
driver.quit()
A selector such as main[data-loaded='true'] is only an example. Wait for the element, text, attribute or application flag that proves your own asynchronous work is complete. readyState == 'complete' covers document loading, not necessarily API calls made afterward.
Command-line alternative
Chrome’s headless command-line mode can serialize the DOM without Selenium:
google-chrome --headless --dump-dom --timeout=10000
--virtual-time-budget=5000
https://example.com/app > rendered.html
--timeout bounds waiting and --virtual-time-budget advances time-dependent JavaScript. Neither option guarantees that a site’s own data-loading condition has completed. The output still references external resources instead of embedding them.
Rank #2
Package the page as one MHTML file
MHTML is the simplest approach when the goal is a portable snapshot rather than a directory of independently named files. The DevTools Protocol command Page.captureSnapshot returns MHTML and documents inclusion of frames, shadow DOM and external resources. It is a DevTools Protocol call, not a standard WebDriver method.
from pathlib import Path
from selenium.webdriver.support.ui import WebDriverWait
url = "https://example.com/app"
driver = make_driver()
try:
driver.get(url)
WebDriverWait(driver, 30).until(
lambda d: d.execute_script("return document.readyState") == "complete"
)
# Add a page-specific readiness wait here when required.
snapshot = driver.execute_cdp_cmd("Page.captureSnapshot", {
"format": "mhtml"
})
Path("page.mhtml").write_text(snapshot["data"], encoding="utf-8")
finally:
driver.quit()
DevTools Protocol’s tip-of-tree definition changes frequently and has no backwards-compatibility guarantee. Pin your implementation to the protocol exposed by the Chrome version you deploy, and treat a protocol error as a version-compatibility issue rather than a Selenium syntax problem.
Extension API option
The Chrome pageCapture extension API can save a tab as MHTML. An extension using it needs the pageCapture permission, and the API is available from Chrome 116. This route is useful when your automation already runs a managed extension; for a standalone Selenium script, the CDP call avoids extension packaging.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteCollect resources individually with network events
Use network capture when you need each response as a separate file, want to inspect headers and status codes, or must reproduce the page’s request graph. Enable tracking before navigation. With CDP, request and response events contain request IDs; retrieve a response body while Chrome still retains it.
import base64
from pathlib import Path
from selenium.webdriver.support.ui import WebDriverWait
url = "https://example.com/app"
out_dir = Path("responses")
out_dir.mkdir(exist_ok=True)
driver = make_driver(performance_logs=False)
try:
driver.execute_cdp_cmd("Network.enable", {})
driver.get(url)
WebDriverWait(driver, 30).until(
lambda d: d.execute_script("return document.readyState") == "complete"
)
# Ask Chrome for events exposed through the performance log only when
# performance logging was enabled at session creation.
events = driver.get_log("performance")
for entry in events:
# Parse entry["message"] and select Network.responseReceived events.
# Then call Network.getResponseBody with each requestId while valid.
pass
finally:
driver.quit()
The abbreviated loop is intentional: production collectors must define safe filenames, preserve content types, follow redirects, cope with duplicate URLs, and decode bodies marked as base64. A complete CDP implementation normally reads Network.responseReceived, stores the requestId and URL, then calls Network.getResponseBody. Some bodies become unavailable after navigation or cache eviction, so retrieve them promptly.
ChromeDriver performance logs
Performance logging is disabled unless requested when the session is created. The capability shown in make_driver(performance_logs=True) enables Network and Page events in ChromeDriver’s performance log. Read the log during the session rather than waiting until after quitting the browser. Performance logs provide event metadata; resource naming, body retrieval and persistence remain your responsibility.
Handle normal browser downloads
A link that triggers a download is different from saving the DOM. Configure Chrome’s download directory, click the element, and wait until temporary files disappear and the expected file exists with a stable size.
Free tools Windows power users keep installed
One-click scans. No signup required.
import time
from pathlib import Path
from selenium.webdriver.support.ui import WebDriverWait
folder = Path("downloads").resolve()
folder.mkdir(exist_ok=True)
driver = make_driver(download_dir=folder)
try:
driver.get("https://example.com/report")
driver.find_element("css selector", "a.download-report").click()
def finished(_):
temporary = list(folder.glob("*.crdownload"))
files = [p for p in folder.iterdir() if p.is_file() and not p.name.endswith(".crdownload")]
return files[0] if files and not temporary else False
downloaded = WebDriverWait(driver, 120, poll_frequency=1).until(finished)
print(f"Saved {downloaded}")
finally:
driver.quit()
ChromeDriver does not wait for downloads automatically. Calling quit() immediately after the click can terminate Chrome before the file is complete. For deterministic jobs, record the expected filename, ignore stale files from earlier runs, and optionally require the size to remain unchanged across two checks.
Waiting, completeness and performance
- Wait on evidence, not elapsed time: prefer a selector, text value, network-idle rule implemented by your application, or a JavaScript readiness flag over a fixed sleep.
- Use a realistic viewport: responsive layouts can load different markup and resources at different widths.
- Expect lazy loading: scroll to relevant sections before capture if images are loaded only when they approach the viewport.
- Bound every wait: a timeout prevents one broken request from holding a worker forever; save diagnostics before aborting.
- Capture logs on failure: keep the current URL, browser console messages, performance events and a screenshot to distinguish a blank page from a selector mismatch.
- Control storage: response archives can be much larger than the HTML. Stream or compress large bodies and sanitize URL-derived filenames.
Troubleshooting
The file contains old or missing content
You probably captured before the application’s asynchronous render finished. Add a page-specific wait and verify the resulting DOM contains a known marker before writing it.
HTML opens but images and styles are absent
A DOM dump stores markup only. Use MHTML for a packaged snapshot or collect network responses separately. Check that requests were not blocked and that the page did not require authentication cookies.
Rank #4
Page.captureSnapshot returns an unknown-command error
The deployed Chrome may not expose that CDP method, or the driver is speaking to a different browser than expected. Check the actual Chrome version and use a protocol definition compatible with it. Do not assume tip-of-tree documentation is stable.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Network events are empty
Enable Network tracking or performance logging before get(url). Performance logging is an opt-in session capability; enabling it after the browser starts is too late for earlier requests.
The download is truncated or missing
Use a dedicated absolute directory, wait for the temporary download suffix to disappear, and only then call quit(). Confirm the Chrome user has write permission.
Chrome will not start
Check Chrome/ChromeDriver compatibility, especially after an automatic browser update. In containers, verify shared-memory and sandbox settings, and make sure the executable is present for the account running Selenium.
The page is blank or blocked
Inspect browser and network logs. A bot check, failed request, authentication redirect or JavaScript exception can produce a valid but useless capture. Treat the page’s own readiness and access requirements as part of the capture design.
Best Value
Or skip the browser setup
For an API-driven screenshot rather than an HTML archive, ScreenshotNeo provides a single GET request and an MCP server for AI clients. It removes cookie banners, newsletter popups and chat widgets before capture; bot checks, blank pages and failed loads are not billed; and every response identifies its page verdict and billing status. The MCP tools take_screenshot, get_page_info and capture_pdf work with Claude, Cursor and other MCP clients. A free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots.
See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, selectors, waits, custom CSS/JavaScript, cookies, headers, device presets, PDFs, caching, signed links and bulk capture.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Create a free ScreenshotNeo account to use the 1,000 monthly screenshots without a card.
FAQ
Is rendered HTML the same as page source?
No. Rendered HTML is serialized after Chrome parses the response and scripts modify the DOM; page source is the original response representation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Should I choose MHTML or separate resources?
Choose MHTML for a convenient one-file snapshot. Choose separate responses when you need to inspect, transform or independently serve each asset.
Can ChromeDriver guarantee that every resource was saved?
No. Readiness, blocked requests, redirects, cache behavior and protocol retention all affect completeness. Define and verify the conditions that matter for your page.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




