Use Puppeteer’s page.content() after navigation and all required interactions have finished. It returns a promise containing the current document’s complete serialized HTML, including the DOCTYPE. For a JavaScript-rendered page, wait for an application-specific signal—such as a selector or item count—before reading it. If you need Chrome’s original, pre-JavaScript source instead, read the HTTPResponse returned by page.goto().
The canonical way to read the current page HTML
This runnable ES module opens a page, waits for network activity to settle, extracts the browser’s current document, prints it, and closes Chromium:
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle2' });
const html = await page.content();
console.log(html);
await browser.close();
The documented signature is content(): Promise<string>. The returned string represents the document as it exists in the page at the moment you call the method, including the DOCTYPE. It is not a frozen copy of the first response: scripts that have already run can insert elements, change attributes, or replace text, and those changes are reflected.
Wait for the state you actually need
networkidle2 is a useful broad navigation condition, but it cannot know when an application has finished rendering its meaningful content. Single-page apps may keep connections open, render after an API response, or require a click. Use a targeted readiness condition whenever possible.
#1 Best Overall
Wait for a root element
await page.goto('https://example.com/app', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#app');
const html = await page.content();
Wait for a known amount of content
await page.goto('https://example.com/products', { waitUntil: 'domcontentloaded' });
await page.waitForFunction(() => {
return document.querySelectorAll('.item').length >= 20;
});
const html = await page.content();
Include an interaction such as “Load more”
await page.goto('https://example.com/products', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#app');
await page.click('.load-more');
await page.waitForFunction(() =>
document.querySelectorAll('.item').length >= 20
);
const html = await page.content();
Prefer a selector, count, or application-specific completion signal to an arbitrary timeout. A fixed sleep can finish too early on a slow run and waste time on a fast one.
Save the extracted HTML correctly
Write the string with an explicit UTF-8 encoding so non-ASCII text is preserved:
import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle2' });
const html = await page.content();
await writeFile('page.html', html, 'utf8');
} finally {
await browser.close();
}
Keeping the browser close in a finally block prevents orphaned Chromium processes when navigation or extraction fails.
page.content() versus outerHTML
If you want to state explicitly that you are serializing the live <html> element in the page context, use evaluate():
const html = await page.evaluate(() => document.documentElement.outerHTML);
Puppeteer executes the function inside the page and returns its result. outerHTML serializes the current <html> element. Formatting can differ from page.content(), while both generally describe the same current DOM. Use page.content() as the normal whole-document API; use evaluate() when you also need page-context logic.
Rank #2
Rendered DOM is not the original “View Source” response
There are two different meanings of “complete source.” Choose deliberately:
| Artifact | How to obtain it | What it contains |
|---|---|---|
| Rendered DOM HTML | await page.content() or document.documentElement.outerHTML |
The document after scripts and interactions have modified it. |
| Original HTTP response body | Read the response returned by page.goto() |
Bytes delivered by the server before browser JavaScript changes the DOM. |
Capture the original response body
const response = await page.goto('https://example.com', {
waitUntil: 'domcontentloaded'
});
if (!response) throw new Error('No main-document response');
const rawSource = await response.text();
console.log(rawSource);
This is the closest match to a browser’s View Source for the main document. It will not include nodes created later by client-side JavaScript. A redirect or unusual navigation can leave you without a main-document response, so check for null before reading the body.
Extract HTML from iframes
An iframe owns a separate document. Serializing the top-level page does not merge that document’s markup into the parent HTML. Find the iframe element, obtain its frame, and call the frame API:
Recommended Free Tools
const iframeElement = await page.waitForSelector('iframe');
const frame = await iframeElement.contentFrame();
if (!frame) throw new Error('iframe frame unavailable');
const iframeHtml = await frame.content();
await writeFile('iframe.html', iframeHtml, 'utf8');
Frame.content() returns the full HTML contents of that frame, including its DOCTYPE. For nested frames, inspect page.frames() or the parent frame’s child frames, select the target, and apply the same method. Keep frame extraction separate from assumptions about the parent document.
Cross-origin frame boundaries
Puppeteer can operate through a frame object when the browser exposes that frame. However, JavaScript running in the page still obeys browser-origin rules. Do not expect a script in the parent page to read arbitrary cross-origin frame variables; extract through the frame object and handle failures independently.
Shadow DOM and what “complete” cannot guarantee
Ordinary document serialization does not expose the internals of a closed shadow root. Even an open component may require explicit traversal if you need its internal markup. Extract an open root in page context:
const componentHtml = await page.evaluate(() => {
const host = document.querySelector('my-component');
if (!host || !host.shadowRoot) return null;
return host.shadowRoot.innerHTML;
});
if (componentHtml === null) {
throw new Error('Open shadow root was not found');
}
Closed roots are intentionally inaccessible through normal page JavaScript. For those components, use the component’s documented API or another supported export rather than assuming the root can be serialized.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchA practical extraction procedure
- Define the artifact. Decide whether you need rendered DOM, the original response, an iframe document, or shadow-root content.
- Navigate. Choose
domcontentloaded,load, ornetworkidle2as a starting condition. - Wait for application readiness. Use
waitForSelector,waitForFunction, or a real application signal. - Perform required interactions. Click controls, expand sections, or paginate before extraction.
- Serialize. Call
page.content(),frame.content(),outerHTML, orresponse.text()according to step one. - Persist and close. Write UTF-8 HTML and close the browser in a
finallyblock.
Performance, reliability, and cost considerations
- Browser startup is expensive. Reuse one browser process and create or close pages per job instead of launching Chromium for every URL.
- Readiness controls correctness. A short navigation wait can produce incomplete markup; an indefinite wait can stall a worker. Set sensible navigation and selector timeouts and fail clearly when the expected signal never appears.
- Large pages consume memory. Full-page DOM strings, embedded data, and many simultaneous tabs increase memory pressure. Process pages in bounded batches and write results rather than retaining every string.
- Network-idle is not universal. Analytics, web sockets, and polling can prevent an idle condition. Prefer a stable selector or application event for those sites.
- Responses and DOM differ by design. Store both when debugging server rendering: the response identifies what was delivered, while the DOM shows what the browser produced.
Troubleshooting common failures
The HTML is missing content loaded by JavaScript
Cause: extraction ran before the app rendered or before an interaction completed. Fix: wait for a specific selector or count, then call page.content(). Replace a guessed delay with waitForFunction tied to the actual state.
networkidle2 never resolves
Cause: persistent requests, polling, or sockets keep network activity alive. Fix: navigate with domcontentloaded and wait for the page’s readiness selector or a bounded application condition.
page.content() does not match View Source
Cause: you compared rendered DOM with the original response. Fix: capture the HTTPResponse from page.goto() and call response.text() for pre-JavaScript source.
Rank #4
The iframe HTML is empty or unavailable
Cause: the frame has not appeared, contentFrame() returned no frame, or the wrong frame was selected. Fix: wait for the iframe element, check the returned frame, and inspect page.frames() for nested targets.
Shadow-root markup is absent
Cause: the component uses a closed shadow root or the markup is inside an open root that ordinary serialization does not include. Fix: query the open shadowRoot in evaluate(); for closed roots, use the component’s supported API.
The process hangs or leaves Chromium running
Cause: an exception bypassed cleanup. Fix: put extraction in try/finally and always call browser.close(). Also bound waits so a missing selector becomes an actionable error.
Or skip the browser setup
If your goal is a clean capture rather than DOM inspection, ScreenshotNeo provides a website screenshot API and MCP server. One request returns PNG, JPEG, WebP, or PDF; it is not a replacement for extracting HTML, but it avoids maintaining a Puppeteer browser when an image or PDF is the required artifact.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options and response details.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesBest Value
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server offers
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan.
Start with 1,000 free screenshots a month—no card required.
Choosing the right extraction method
| Your requirement | Use |
|---|---|
| Current, post-JavaScript document | page.content() |
| Explicit live DOM serialization | page.evaluate(() => document.documentElement.outerHTML) |
| Main document before scripts modify it | response.text() from page.goto() |
| Iframe document | frame.content() |
| Open shadow-root internals | Page-context access through shadowRoot |
Frequently Asked Questions
Does page.content() include the DOCTYPE?
Yes. Puppeteer documents it as returning the full HTML contents of the page, including the DOCTYPE.
Can Puppeteer return the source of a page that redirects?
It can return the final page’s rendered DOM. For original-response capture, verify the response returned by the navigation and handle a missing main-document response explicitly.
Should I use outerHTML or page.content()?
Use page.content() for the normal whole-document operation. Use outerHTML when you need page-context logic or want to serialize the live <html> element directly.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →The Bottom Line
Call page.content() only after the exact rendered state you need is ready; use response.text() for original server source, and extract iframe or open shadow-root documents separately.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




