Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Puppeteer to let the page’s JavaScript run, wait for the specific element or value you need, then extract and validate it in the browser context. A targeted wait such as waitForSelector or waitForFunction is usually more reliable than sleeping for an arbitrary number of seconds.

Why the initial HTML may not contain the value

A website can send an initial HTML document that contains little more than a shell. Client-side JavaScript may then request data, calculate a value, and insert it into the DOM. Reading the original response alone can therefore miss content that appears later. Puppeteer controls a real browser context, so the page’s JavaScript can execute before you inspect the rendered page. Puppeteer’s documentation describes it as a JavaScript library for controlling Chrome or Firefox over the DevTools Protocol or WebDriver BiDi: Puppeteer documentation.

The important distinction is between page navigation and application readiness. A navigation event tells you something about document loading; it does not prove that the particular price, status, or other value you want has been rendered.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A reliable scrape: wait for the data, then extract it

  1. Open a browser page. Launch Puppeteer or connect to a browser, then create a page.
  2. Navigate to the target URL. Choose an appropriate navigation condition, such as domcontentloaded.
  3. Identify where the value lives. It may be an element’s text, an attribute, or a value represented in the page’s DOM.
  4. Wait for that condition. Use waitForSelector for an element, or waitForFunction if an existing element changes later.
  5. Extract and validate. Read the text or attribute, then confirm it is non-empty and has the expected format before saving it.

This example waits for a visible element with a data-price attribute and reads its trimmed text. Replace the URL and selector with ones that match the site you are permitted to access.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });

try {
  const page = await browser.newPage();
  await page.goto('https://example.com/product', {
    waitUntil: 'domcontentloaded',
  });

  await page.waitForSelector('[data-price]', {
    visible: true,
    timeout: 15000,
  });

  const price = await page.$eval(
    '[data-price]',
    el => el.textContent?.trim() ?? ''
  );

  if (!price) {
    throw new Error('The price element was present but contained no text');
  }

  console.log(price);
} finally {
  await browser.close();
}

The try/finally ensures the browser is closed if navigation or extraction fails. Set a timeout appropriate to the site rather than disabling timeouts by default.

Choose the right readiness condition

Wait strategy Use it when Trade-off
waitForSelector The target element is inserted into the DOM after JavaScript runs. It confirms a matching element exists; it does not by itself guarantee the element’s text or attribute has reached its final value.
waitForFunction The element exists early, but its text, attribute, or state changes later. The predicate must describe the actual ready state. A predicate that is too broad can resolve before useful data appears.
waitForNetworkIdle You need network quiescence as an additional synchronization signal. It measures network activity, not application readiness. Polling, analytics, WebSockets, or lazy loading can delay it, while a page may become idle before the target value is ready.
Fixed delay Only as a last-resort workaround when there is no observable readiness condition. A guessed delay can be wasteful on fast loads and too short on slow ones; it does not make extraction deterministic.

Wait for an element with waitForSelector

Use this when the site inserts a distinct node when the data is ready. The current API accepts visibility options, and supports a timeout or abort signal. Its documented default timeout is 30 seconds; timeout: 0 disables the timeout. If the selector does not appear within the timeout, the wait throws rather than silently continuing. See the Page.waitForSelector API reference.

  • visible: true is appropriate when the value must be visible to a user.
  • hidden: true waits for an element to be removed or concealed; it is useful for loading overlays, but does not prove the requested data itself has loaded.
  • Keep the timeout bounded so failures are visible and a stuck page cannot wait indefinitely.

Wait for a value with waitForFunction

If the element appears before its content is populated, wait for a meaningful value rather than the node alone:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await page.waitForFunction(() => {
  const value = document.querySelector('[data-total]')?.textContent?.trim();
  return Boolean(value);
}, { timeout: 15000 });

const total = await page.$eval(
  '[data-total]',
  el => el.textContent?.trim() ?? ''
);

The predicate executes in the page, so it can inspect the rendered DOM. Make the condition specific enough to exclude placeholders such as an empty string or “Loading…”. Page.evaluate runs a function in the page context and waits for a returned Promise to resolve; keep browser-side functions self-contained and pass any needed inputs as arguments rather than relying on Node.js variables. See the Page.evaluate API reference.

Use network-idle waiting as supporting evidence

page.waitForNetworkIdle() waits for network activity to become idle and waits at least for the configured idle time. It can be useful alongside a data-specific wait, but “network idle” is not the same as “the application has finished rendering.” A site that keeps connections open may never reach the state you expect; another may render the value after its relevant request has completed. Prefer a selector or predicate tied to the value you need. See the Page.waitForNetworkIdle API reference.

Extract text, attributes, or a list

Read one element

After waiting, use page.$eval to run a function on the first matching element. Use textContent for text, or getAttribute for an HTML attribute. Normalize and validate the result before treating it as usable data.

const value = await page.$eval(
  '[data-status]',
  el => el.getAttribute('data-status')?.trim() ?? ''
);

if (!value) {
  throw new Error('Status attribute was empty or missing');
}

Read matching elements with $$eval

For repeated rows, page.$$eval passes all matching nodes to a function and returns the function’s result. This example gathers text and an attribute from each row:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const rows = await page.$$eval('[data-row]', nodes =>
  nodes.map(node => ({
    name: node.querySelector('.name')?.textContent?.trim() ?? '',
    value: node.getAttribute('data-value') ?? '',
  }))
);

const usableRows = rows.filter(row => row.name && row.value);

For a single value, $eval is direct; for a collection, $$eval avoids a separate browser evaluation for every row. The Page.$$eval API reference documents the method.

Make selectors resilient to frontend changes

Prefer stable semantic attributes, labels, or roles when a site provides them. A class used only for styling may change during a redesign even though the underlying content is unchanged. Puppeteer’s page-interactions guide covers CSS selectors and additional selector options for text, accessibility attributes, XPath, and shadow-root traversal: Page interactions.

  • Use a data attribute or an accessible role/name if it identifies the data reliably.
  • Scope a selector to a relevant container when the same label appears elsewhere on the page.
  • For values inside an iframe, identify the relevant frame and perform the wait and extraction there.
  • For values inside a shadow root, use a selector strategy that reaches that root; a regular document query may not find the node.

Troubleshoot empty or incorrect results

The selector times out

First confirm that the selector is valid and that the node appears in the rendered page, not only in a different frame or shadow root. The value may require a prior click, scroll, consent action, or pagination step. During debugging, capture a screenshot or inspect await page.content() after the wait fails to see what Puppeteer actually rendered.

The element exists but the result is empty

The node may be a placeholder that gets updated later. Change the readiness condition to test the text or attribute itself with waitForFunction, then read it. Also check that you are reading the right field: a visible label may differ from the machine-readable attribute holding the value.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The value is in an iframe or shadow root

A selector evaluated against the top-level document cannot necessarily reach content in another browsing context or an encapsulated shadow tree. Locate the frame or use a selector strategy that supports the relevant shadow root, then repeat the same wait-and-extract sequence there.

Network-idle waiting hangs or finishes too soon

Long-lived requests, polling, analytics, or WebSockets can prevent network quiescence. Conversely, a quiet network does not guarantee that delayed frontend code has updated the page. Use a data-specific selector or predicate as the decisive condition, with network-idle only as an optional supporting wait.

The browser waits too long or fails unpredictably

Set explicit, bounded timeouts on waits and handle their errors so a missing value is reported as a failed scrape, not saved as valid output. Avoid disabling timeouts unless the surrounding job has its own cancellation and recovery policy. If navigation itself fails, report that separately from a selector timeout; they indicate different points of failure.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and responsible collection

Waiting for the exact value often improves both speed and reliability compared with a long fixed delay: a ready page can proceed without waiting out an unnecessary sleep, while a slower page does not get scraped prematurely. There is no universal wait duration or performance number for every site; choose timeouts for your target and make timeout failures observable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep browser resources under control by closing pages and browsers when the job finishes, including on errors. Validate output before writing it to a database or file, and distinguish navigation failure, readiness timeout, and invalid extracted data in logs. For repeated collection, use the site’s authorized access methods and respect its terms and applicable law.

Best Value
The SQL Programming Language: .
  • Used Book in Good Condition

Or skip the browser setup

If you need a screenshot rather than structured DOM data, ScreenshotNeo can return a website capture with one GET request. Its clean-shot steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, with the response identifying the page verdict and billing status in headers. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents and other MCP clients. Read the ScreenshotNeo API documentation for request options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://example.com/product 
  -o shot.webp

The free plan includes 1,000 screenshots per month with no card required; paid plans start at $5 for 3,000 shots. Visit ScreenshotNeo to learn more, or sign up free for 1,000 screenshots a month with no card.

Frequently Asked Questions

Can Puppeteer extract a value that is not visible on screen?

Yes. If the value exists in the rendered DOM or an attribute, extract it without requiring visibility; use visible: true only when visibility is part of the condition you need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use page.content() to scrape the rendered value?

Use it to inspect or debug the rendered HTML. For routine extraction, a targeted $eval or $$eval avoids parsing the whole document.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.