Use await page.content() after navigation and any condition that means the page is ready. It returns the browser page’s complete HTML, including the DOCTYPE:
const response = await page.goto('https://example.com');
if (!response) throw new Error('Navigation did not produce a response');
await page.waitForSelector('main');
const html = await page.content();
console.log(html);
This is the page as Puppeteer’s browser currently represents it—not necessarily the original response bytes sent by the server. The right method depends on whether you need the whole document, one element, an iframe, or the untouched HTTP response.
What “page source” means in Puppeteer
In a normal browser, “view source” usually means the HTML response received from the server. In Puppeteer, developers often mean the post-render DOM: the markup after scripts have inserted content, changed attributes, or removed nodes. Those are different artifacts.
- Rendered document: the current page DOM, retrieved with
page.content()or DOM evaluation. - Selected markup: the HTML inside a particular element, retrieved with
$eval()or an evaluation function. - Original response: the navigation response body captured separately before or while the browser parses it.
Choose the representation before writing your scraper or test. A serialized DOM is useful for what a user can see after JavaScript runs; it is not a byte-for-byte archive of the server response.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Get the complete current HTML with page.content()
Page.content() is the direct answer for a page’s full current HTML. Puppeteer documents it as returning the full HTML contents, including the DOCTYPE, as a Promise<string>.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({headless: true});
try {
const page = await browser.newPage();
const response = await page.goto('https://example.com', {
waitUntil: 'domcontentloaded'
});
if (!response) throw new Error('No navigation response');
const html = await page.content();
console.log(html);
} finally {
await browser.close();
}
waitUntil: 'domcontentloaded' tells Puppeteer to wait for the initial document parse. It does not guarantee that a single-page application has finished fetching and rendering its data. Add an application-specific readiness condition when later content matters.
Wait for the content you actually need
Do not replace a meaningful readiness signal with an arbitrary fixed sleep. Use the condition that identifies the content your program must capture.
Wait for a selector
await page.goto('https://example.com/dashboard', {
waitUntil: 'domcontentloaded'
});
await page.waitForSelector('[data-ready="true"]');
const html = await page.content();
The selector should represent a real application state: a results container, a table row, a page-specific marker, or another element that appears only when the required render has completed.
Wait for a function
await page.goto(url, {waitUntil: 'domcontentloaded'});
await page.waitForFunction(() => {
const status = document.querySelector('#status');
return status?.textContent?.trim() === 'Loaded';
});
const html = await page.content();
This is useful when readiness depends on text, a count, a class, or another browser-side condition rather than simple element existence.
Wait for network idle
await page.goto(url, {waitUntil: 'domcontentloaded'});
await page.waitForNetworkIdle({idleTime: 500, timeout: 30000});
const html = await page.content();
Network idle can help with pages that finish rendering after several requests, but it is not universal: analytics, polling, WebSockets, or advertisements can keep traffic active. Prefer a page-specific signal when one exists.
Serialize the DOM with evaluate()
page.evaluate() runs JavaScript in the page context and returns the value produced by that function. To serialize the document element explicitly:
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
const html = await page.evaluate(() =>
document.documentElement.outerHTML
);
console.log(html);
For the body only:
const bodyHtml = await page.evaluate(() => document.body.innerHTML);
For most whole-document retrieval, page.content() is clearer. Evaluation is useful when you need a deliberately chosen DOM property or want to transform the value before it crosses back to Node.js.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Read one element instead of the entire document
When a scraper needs only a region, use page.$eval(). Puppeteer passes the first matching element to your function and throws if no element matches, so make the selector and error handling explicit.
const mainHtml = await page.$eval(
'main',
element => element.innerHTML
);
console.log(mainHtml);
This returns the selected element’s contents, not the element’s own opening and closing tags. To include the element itself:
const mainWithTag = await page.$eval(
'main',
element => element.outerHTML
);
A missing selector is normally a timing problem, a selector problem, or a page variation. Wait for it when rendering is asynchronous and verify the selector against the current page.
Page source versus the original HTTP response
page.content(), outerHTML, and innerHTML describe the browser’s current DOM. Browser parsing can normalize markup, scripts can mutate it, and dynamically fetched content may not exist in the initial response.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteIf you need the original response body, capture the navigation response separately:
const response = await page.goto(url, {waitUntil: 'domcontentloaded'});
if (!response) throw new Error('Navigation did not produce a response');
const responseText = await response.text();
const renderedHtml = await page.content();
Keep the two values separate in your data model. The response text is appropriate for byte-oriented or server-source analysis; the rendered HTML is appropriate for inspecting the document after browser execution. Also inspect the response status when HTTP success matters. Valid HTTP responses such as 404 or 500 do not necessarily cause goto() to throw.
Rank #3
Get markup from an iframe
An iframe has its own document and execution context. The top-level page’s content() does not automatically give you the child document’s HTML.
await page.goto('https://example.com', {waitUntil: 'domcontentloaded'});
await page.waitForSelector('iframe');
const frame = page.frames().find(f =>
f.url().includes('/embedded-content')
);
if (!frame) throw new Error('Target frame not found');
await frame.waitForSelector('main');
const frameHtml = await frame.content();
console.log(frameHtml);
If the frame is same-origin and you only need a selected region, evaluate in that frame:
const frameMain = await frame.$eval(
'main',
element => element.innerHTML
);
Use the frame’s URL, name, or another stable property to identify it. Do not assume the first attached frame is always the one you want; pages commonly contain analytics, payment, or advertising frames.
Common failures and precise fixes
The HTML is missing JavaScript-rendered content
Cause: navigation completed before the application finished rendering.
Fix: wait for the result container, a readiness attribute, a specific function condition, or an appropriate network-idle state, then call content(). Avoid a guessed sleep when a reliable signal exists.
content() returns more than the body
Cause: this is expected behavior. The method returns the full document, including the DOCTYPE.
Fix: use document.body.innerHTML or $eval() for a narrower result.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
setContent() was used as if it were a getter
Cause: page.setContent(html) assigns markup to a page and returns a promise for completion; it does not retrieve source.
Fix: call await page.content() after setting the content if you need to read it back.
await page.setContent('<main><h1>Hello</h1></main>');
const html = await page.content();
$eval() throws “failed to find element”
Cause: the selector did not match at the moment of evaluation.
Free tools Windows power users keep installed
One-click scans. No signup required.
Fix: check the selector, wait for it, and account for responsive or logged-in variants:
await page.waitForSelector('main', {timeout: 15000});
const html = await page.$eval('main', el => el.outerHTML);
The wrong document was captured from an iframe
Cause: the top-level page and each child frame have separate documents.
Fix: locate the intended frame with page.frames(), wait in that frame, and call frame.content() or frame.$eval().
The status is an error but navigation did not throw
Cause: navigation can resolve for valid HTTP error responses.
Recommended Free Tools
Best Value
Fix: inspect response.status() and decide whether to reject non-success statuses:
const response = await page.goto(url);
if (!response) throw new Error('No response');
if (response.status() >= 400) {
throw new Error(`HTTP ${response.status()}`);
}
The result is not the server’s exact source
Cause: DOM serialization reflects browser parsing and script changes.
Fix: capture response.text() for the navigation response, or use a direct HTTP client when browser execution is not required.
Choosing the right Puppeteer method
| Need | Method | Result | Important timing note |
|---|---|---|---|
| Whole current document | page.content() |
HTML including DOCTYPE | Call after the required render state |
| Explicit DOM serialization | page.evaluate(() => document.documentElement.outerHTML) |
Document element’s serialized HTML | Uses the current browser DOM |
| One region | page.$eval(selector, el => el.innerHTML) |
Contents of the first match | Throws if no element matches |
| Child document | frame.content() |
HTML for the selected iframe document | Wait in the frame’s own context |
| Markup assignment | page.setContent(html) |
Writes HTML to the page | Read afterward with content() |
| Original response body | Navigation response.text() |
Response text captured separately | Not equivalent to rendered DOM |
Or skip the browser setup
If your goal is a clean screenshot or PDF rather than HTML inspection, ScreenshotNeo provides a single-call website capture API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For a screenshot, use the API shown in the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
It also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDFs with paper size and page ranges, HTML/CSS rendering, custom JavaScript and CSS, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work for easier migration.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to start.
Performance and reliability considerations
- Capture only what you need: use a selected element when a full document would create unnecessary memory and parsing work.
- Reuse a browser: launch Puppeteer once for a batch, create pages as needed, and close pages when each job finishes.
- Bound waits: set practical timeouts on selectors and functions so a broken application state cannot hang a worker indefinitely.
- Record context: save the URL, timestamp, status code, readiness condition, and whether the result came from the response or rendered DOM.
- Handle authentication and variants: source can differ by cookies, viewport, locale, user agent, and logged-in state; make those inputs explicit.
- Protect sensitive HTML: rendered source can contain personal data, tokens in attributes, or private application state. Store and log it accordingly.
Recommended decision path
- Decide whether you need the original response or the post-JavaScript DOM.
- For the complete rendered document, navigate, wait for the page-specific condition, and call
page.content(). - For a deliberate DOM serialization, use
evaluate(); for one region, use$eval(). - For an iframe, locate its
Frameand run the same operation in that context. - Inspect navigation status separately when HTTP success is required.
- Keep response text and rendered HTML as separate outputs when both forms matter.
Frequently Asked Questions
Does Puppeteer have a method named `getPageSource()`?
No. Use `page.content()` for the complete current document, or the narrower DOM methods when that is what your task requires.
Will `page.content()` execute JavaScript?
JavaScript executes as part of normal page loading in the browser; `content()` then serializes the DOM state that exists when you call it. It does not itself wait for an application-specific render condition.
Can I use this for XML?
Only if the browser page exposes the XML document in the context you inspect. For exact response bytes or XML-specific parsing, capture the navigation response or use an HTTP client instead of relying on DOM serialization.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

