Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use await page.content() after navigation and any condition that means the page is ready. It returns the browser page’s complete HTML, including the DOCTYPE:

const response = await page.goto('https://example.com');
if (!response) throw new Error('Navigation did not produce a response');
await page.waitForSelector('main');
const html = await page.content();
console.log(html);

This is the page as Puppeteer’s browser currently represents it—not necessarily the original response bytes sent by the server. The right method depends on whether you need the whole document, one element, an iframe, or the untouched HTTP response.

What “page source” means in Puppeteer

In a normal browser, “view source” usually means the HTML response received from the server. In Puppeteer, developers often mean the post-render DOM: the markup after scripts have inserted content, changed attributes, or removed nodes. Those are different artifacts.

  • Rendered document: the current page DOM, retrieved with page.content() or DOM evaluation.
  • Selected markup: the HTML inside a particular element, retrieved with $eval() or an evaluation function.
  • Original response: the navigation response body captured separately before or while the browser parses it.

Choose the representation before writing your scraper or test. A serialized DOM is useful for what a user can see after JavaScript runs; it is not a byte-for-byte archive of the server response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Get the complete current HTML with page.content()

Page.content() is the direct answer for a page’s full current HTML. Puppeteer documents it as returning the full HTML contents, including the DOCTYPE, as a Promise<string>.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({headless: true});
try {
  const page = await browser.newPage();
  const response = await page.goto('https://example.com', {
    waitUntil: 'domcontentloaded'
  });

  if (!response) throw new Error('No navigation response');
  const html = await page.content();
  console.log(html);
} finally {
  await browser.close();
}

waitUntil: 'domcontentloaded' tells Puppeteer to wait for the initial document parse. It does not guarantee that a single-page application has finished fetching and rendering its data. Add an application-specific readiness condition when later content matters.

Wait for the content you actually need

Do not replace a meaningful readiness signal with an arbitrary fixed sleep. Use the condition that identifies the content your program must capture.

Wait for a selector

await page.goto('https://example.com/dashboard', {
  waitUntil: 'domcontentloaded'
});
await page.waitForSelector('[data-ready="true"]');
const html = await page.content();

The selector should represent a real application state: a results container, a table row, a page-specific marker, or another element that appears only when the required render has completed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for a function

await page.goto(url, {waitUntil: 'domcontentloaded'});
await page.waitForFunction(() => {
  const status = document.querySelector('#status');
  return status?.textContent?.trim() === 'Loaded';
});
const html = await page.content();

This is useful when readiness depends on text, a count, a class, or another browser-side condition rather than simple element existence.

Wait for network idle

await page.goto(url, {waitUntil: 'domcontentloaded'});
await page.waitForNetworkIdle({idleTime: 500, timeout: 30000});
const html = await page.content();

Network idle can help with pages that finish rendering after several requests, but it is not universal: analytics, polling, WebSockets, or advertisements can keep traffic active. Prefer a page-specific signal when one exists.

Serialize the DOM with evaluate()

page.evaluate() runs JavaScript in the page context and returns the value produced by that function. To serialize the document element explicitly:

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option
const html = await page.evaluate(() =>
  document.documentElement.outerHTML
);
console.log(html);

For the body only:

const bodyHtml = await page.evaluate(() => document.body.innerHTML);

For most whole-document retrieval, page.content() is clearer. Evaluation is useful when you need a deliberately chosen DOM property or want to transform the value before it crosses back to Node.js.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read one element instead of the entire document

When a scraper needs only a region, use page.$eval(). Puppeteer passes the first matching element to your function and throws if no element matches, so make the selector and error handling explicit.

const mainHtml = await page.$eval(
  'main',
  element => element.innerHTML
);
console.log(mainHtml);

This returns the selected element’s contents, not the element’s own opening and closing tags. To include the element itself:

const mainWithTag = await page.$eval(
  'main',
  element => element.outerHTML
);

A missing selector is normally a timing problem, a selector problem, or a page variation. Wait for it when rendering is asynchronous and verify the selector against the current page.

Page source versus the original HTTP response

page.content(), outerHTML, and innerHTML describe the browser’s current DOM. Browser parsing can normalize markup, scripts can mutate it, and dynamically fetched content may not exist in the initial response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If you need the original response body, capture the navigation response separately:

const response = await page.goto(url, {waitUntil: 'domcontentloaded'});
if (!response) throw new Error('Navigation did not produce a response');
const responseText = await response.text();
const renderedHtml = await page.content();

Keep the two values separate in your data model. The response text is appropriate for byte-oriented or server-source analysis; the rendered HTML is appropriate for inspecting the document after browser execution. Also inspect the response status when HTTP success matters. Valid HTTP responses such as 404 or 500 do not necessarily cause goto() to throw.

Get markup from an iframe

An iframe has its own document and execution context. The top-level page’s content() does not automatically give you the child document’s HTML.

await page.goto('https://example.com', {waitUntil: 'domcontentloaded'});
await page.waitForSelector('iframe');

const frame = page.frames().find(f =>
  f.url().includes('/embedded-content')
);
if (!frame) throw new Error('Target frame not found');

await frame.waitForSelector('main');
const frameHtml = await frame.content();
console.log(frameHtml);

If the frame is same-origin and you only need a selected region, evaluate in that frame:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const frameMain = await frame.$eval(
  'main',
  element => element.innerHTML
);

Use the frame’s URL, name, or another stable property to identify it. Do not assume the first attached frame is always the one you want; pages commonly contain analytics, payment, or advertising frames.

Common failures and precise fixes

The HTML is missing JavaScript-rendered content

Cause: navigation completed before the application finished rendering.

Fix: wait for the result container, a readiness attribute, a specific function condition, or an appropriate network-idle state, then call content(). Avoid a guessed sleep when a reliable signal exists.

content() returns more than the body

Cause: this is expected behavior. The method returns the full document, including the DOCTYPE.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fix: use document.body.innerHTML or $eval() for a narrower result.

Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

setContent() was used as if it were a getter

Cause: page.setContent(html) assigns markup to a page and returns a promise for completion; it does not retrieve source.

Fix: call await page.content() after setting the content if you need to read it back.

await page.setContent('<main><h1>Hello</h1></main>');
const html = await page.content();

$eval() throws “failed to find element”

Cause: the selector did not match at the moment of evaluation.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fix: check the selector, wait for it, and account for responsive or logged-in variants:

await page.waitForSelector('main', {timeout: 15000});
const html = await page.$eval('main', el => el.outerHTML);

The wrong document was captured from an iframe

Cause: the top-level page and each child frame have separate documents.

Fix: locate the intended frame with page.frames(), wait in that frame, and call frame.content() or frame.$eval().

The status is an error but navigation did not throw

Cause: navigation can resolve for valid HTTP error responses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fix: inspect response.status() and decide whether to reject non-success statuses:

const response = await page.goto(url);
if (!response) throw new Error('No response');
if (response.status() >= 400) {
  throw new Error(`HTTP ${response.status()}`);
}

The result is not the server’s exact source

Cause: DOM serialization reflects browser parsing and script changes.

Fix: capture response.text() for the navigation response, or use a direct HTTP client when browser execution is not required.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choosing the right Puppeteer method

Need Method Result Important timing note
Whole current document page.content() HTML including DOCTYPE Call after the required render state
Explicit DOM serialization page.evaluate(() => document.documentElement.outerHTML) Document element’s serialized HTML Uses the current browser DOM
One region page.$eval(selector, el => el.innerHTML) Contents of the first match Throws if no element matches
Child document frame.content() HTML for the selected iframe document Wait in the frame’s own context
Markup assignment page.setContent(html) Writes HTML to the page Read afterward with content()
Original response body Navigation response.text() Response text captured separately Not equivalent to rendered DOM

Or skip the browser setup

If your goal is a clean screenshot or PDF rather than HTML inspection, ScreenshotNeo provides a single-call website capture API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a screenshot, use the API shown in the ScreenshotNeo documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());

It also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDFs with paper size and page ranges, HTML/CSS rendering, custom JavaScript and CSS, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work for easier migration.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to start.

Performance and reliability considerations

  • Capture only what you need: use a selected element when a full document would create unnecessary memory and parsing work.
  • Reuse a browser: launch Puppeteer once for a batch, create pages as needed, and close pages when each job finishes.
  • Bound waits: set practical timeouts on selectors and functions so a broken application state cannot hang a worker indefinitely.
  • Record context: save the URL, timestamp, status code, readiness condition, and whether the result came from the response or rendered DOM.
  • Handle authentication and variants: source can differ by cookies, viewport, locale, user agent, and logged-in state; make those inputs explicit.
  • Protect sensitive HTML: rendered source can contain personal data, tokens in attributes, or private application state. Store and log it accordingly.

Recommended decision path

  1. Decide whether you need the original response or the post-JavaScript DOM.
  2. For the complete rendered document, navigate, wait for the page-specific condition, and call page.content().
  3. For a deliberate DOM serialization, use evaluate(); for one region, use $eval().
  4. For an iframe, locate its Frame and run the same operation in that context.
  5. Inspect navigation status separately when HTTP success is required.
  6. Keep response text and rendered HTML as separate outputs when both forms matter.

Frequently Asked Questions

Does Puppeteer have a method named `getPageSource()`?

No. Use `page.content()` for the complete current document, or the narrower DOM methods when that is what your task requires.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Will `page.content()` execute JavaScript?

JavaScript executes as part of normal page loading in the browser; `content()` then serializes the DOM state that exists when you call it. It does not itself wait for an application-specific render condition.

Can I use this for XML?

Only if the browser page exposes the XML document in the context you inspect. For exact response bytes or XML-specific parsing, capture the navigation response or use an HTTP client instead of relying on DOM serialization.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.