Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Puppeteer to let the page’s JavaScript run, wait for the specific element or value you need, then extract and validate it in the browser context. A targeted wait such as waitForSelector or waitForFunction is usually more reliable than sleeping for an arbitrary number of seconds.
Why the initial HTML may not contain the value
A website can send an initial HTML document that contains little more than a shell. Client-side JavaScript may then request data, calculate a value, and insert it into the DOM. Reading the original response alone can therefore miss content that appears later. Puppeteer controls a real browser context, so the page’s JavaScript can execute before you inspect the rendered page. Puppeteer’s documentation describes it as a JavaScript library for controlling Chrome or Firefox over the DevTools Protocol or WebDriver BiDi: Puppeteer documentation.
The important distinction is between page navigation and application readiness. A navigation event tells you something about document loading; it does not prove that the particular price, status, or other value you want has been rendered.
A reliable scrape: wait for the data, then extract it
- Open a browser page. Launch Puppeteer or connect to a browser, then create a page.
- Navigate to the target URL. Choose an appropriate navigation condition, such as
domcontentloaded. - Identify where the value lives. It may be an element’s text, an attribute, or a value represented in the page’s DOM.
- Wait for that condition. Use
waitForSelectorfor an element, orwaitForFunctionif an existing element changes later. - Extract and validate. Read the text or attribute, then confirm it is non-empty and has the expected format before saving it.
This example waits for a visible element with a data-price attribute and reads its trimmed text. Replace the URL and selector with ones that match the site you are permitted to access.
#1 Best Overall
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/product', {
waitUntil: 'domcontentloaded',
});
await page.waitForSelector('[data-price]', {
visible: true,
timeout: 15000,
});
const price = await page.$eval(
'[data-price]',
el => el.textContent?.trim() ?? ''
);
if (!price) {
throw new Error('The price element was present but contained no text');
}
console.log(price);
} finally {
await browser.close();
}
The try/finally ensures the browser is closed if navigation or extraction fails. Set a timeout appropriate to the site rather than disabling timeouts by default.
Choose the right readiness condition
| Wait strategy | Use it when | Trade-off |
|---|---|---|
waitForSelector |
The target element is inserted into the DOM after JavaScript runs. | It confirms a matching element exists; it does not by itself guarantee the element’s text or attribute has reached its final value. |
waitForFunction |
The element exists early, but its text, attribute, or state changes later. | The predicate must describe the actual ready state. A predicate that is too broad can resolve before useful data appears. |
waitForNetworkIdle |
You need network quiescence as an additional synchronization signal. | It measures network activity, not application readiness. Polling, analytics, WebSockets, or lazy loading can delay it, while a page may become idle before the target value is ready. |
| Fixed delay | Only as a last-resort workaround when there is no observable readiness condition. | A guessed delay can be wasteful on fast loads and too short on slow ones; it does not make extraction deterministic. |
Wait for an element with waitForSelector
Use this when the site inserts a distinct node when the data is ready. The current API accepts visibility options, and supports a timeout or abort signal. Its documented default timeout is 30 seconds; timeout: 0 disables the timeout. If the selector does not appear within the timeout, the wait throws rather than silently continuing. See the Page.waitForSelector API reference.
visible: trueis appropriate when the value must be visible to a user.hidden: truewaits for an element to be removed or concealed; it is useful for loading overlays, but does not prove the requested data itself has loaded.- Keep the timeout bounded so failures are visible and a stuck page cannot wait indefinitely.
Wait for a value with waitForFunction
If the element appears before its content is populated, wait for a meaningful value rather than the node alone:
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesawait page.waitForFunction(() => {
const value = document.querySelector('[data-total]')?.textContent?.trim();
return Boolean(value);
}, { timeout: 15000 });
const total = await page.$eval(
'[data-total]',
el => el.textContent?.trim() ?? ''
);
The predicate executes in the page, so it can inspect the rendered DOM. Make the condition specific enough to exclude placeholders such as an empty string or “Loading…”. Page.evaluate runs a function in the page context and waits for a returned Promise to resolve; keep browser-side functions self-contained and pass any needed inputs as arguments rather than relying on Node.js variables. See the Page.evaluate API reference.
Rank #2
Use network-idle waiting as supporting evidence
page.waitForNetworkIdle() waits for network activity to become idle and waits at least for the configured idle time. It can be useful alongside a data-specific wait, but “network idle” is not the same as “the application has finished rendering.” A site that keeps connections open may never reach the state you expect; another may render the value after its relevant request has completed. Prefer a selector or predicate tied to the value you need. See the Page.waitForNetworkIdle API reference.
Extract text, attributes, or a list
Read one element
After waiting, use page.$eval to run a function on the first matching element. Use textContent for text, or getAttribute for an HTML attribute. Normalize and validate the result before treating it as usable data.
const value = await page.$eval(
'[data-status]',
el => el.getAttribute('data-status')?.trim() ?? ''
);
if (!value) {
throw new Error('Status attribute was empty or missing');
}
Read matching elements with $$eval
For repeated rows, page.$$eval passes all matching nodes to a function and returns the function’s result. This example gathers text and an attribute from each row:
Free tools Windows power users keep installed
One-click scans. No signup required.
const rows = await page.$$eval('[data-row]', nodes =>
nodes.map(node => ({
name: node.querySelector('.name')?.textContent?.trim() ?? '',
value: node.getAttribute('data-value') ?? '',
}))
);
const usableRows = rows.filter(row => row.name && row.value);
For a single value, $eval is direct; for a collection, $$eval avoids a separate browser evaluation for every row. The Page.$$eval API reference documents the method.
Make selectors resilient to frontend changes
Prefer stable semantic attributes, labels, or roles when a site provides them. A class used only for styling may change during a redesign even though the underlying content is unchanged. Puppeteer’s page-interactions guide covers CSS selectors and additional selector options for text, accessibility attributes, XPath, and shadow-root traversal: Page interactions.
- Use a data attribute or an accessible role/name if it identifies the data reliably.
- Scope a selector to a relevant container when the same label appears elsewhere on the page.
- For values inside an iframe, identify the relevant frame and perform the wait and extraction there.
- For values inside a shadow root, use a selector strategy that reaches that root; a regular document query may not find the node.
Troubleshoot empty or incorrect results
The selector times out
First confirm that the selector is valid and that the node appears in the rendered page, not only in a different frame or shadow root. The value may require a prior click, scroll, consent action, or pagination step. During debugging, capture a screenshot or inspect await page.content() after the wait fails to see what Puppeteer actually rendered.
The element exists but the result is empty
The node may be a placeholder that gets updated later. Change the readiness condition to test the text or attribute itself with waitForFunction, then read it. Also check that you are reading the right field: a visible label may differ from the machine-readable attribute holding the value.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The value is in an iframe or shadow root
A selector evaluated against the top-level document cannot necessarily reach content in another browsing context or an encapsulated shadow tree. Locate the frame or use a selector strategy that supports the relevant shadow root, then repeat the same wait-and-extract sequence there.
Rank #4
Network-idle waiting hangs or finishes too soon
Long-lived requests, polling, analytics, or WebSockets can prevent network quiescence. Conversely, a quiet network does not guarantee that delayed frontend code has updated the page. Use a data-specific selector or predicate as the decisive condition, with network-idle only as an optional supporting wait.
The browser waits too long or fails unpredictably
Set explicit, bounded timeouts on waits and handle their errors so a missing value is reported as a failed scrape, not saved as valid output. Avoid disabling timeouts unless the surrounding job has its own cancellation and recovery policy. If navigation itself fails, report that separately from a selector timeout; they indicate different points of failure.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance, reliability, and responsible collection
Waiting for the exact value often improves both speed and reliability compared with a long fixed delay: a ready page can proceed without waiting out an unnecessary sleep, while a slower page does not get scraped prematurely. There is no universal wait duration or performance number for every site; choose timeouts for your target and make timeout failures observable.
Keep browser resources under control by closing pages and browsers when the job finishes, including on errors. Validate output before writing it to a database or file, and distinguish navigation failure, readiness timeout, and invalid extracted data in logs. For repeated collection, use the site’s authorized access methods and respect its terms and applicable law.
Best Value
- Used Book in Good Condition
Or skip the browser setup
If you need a screenshot rather than structured DOM data, ScreenshotNeo can return a website capture with one GET request. Its clean-shot steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, with the response identifying the page verdict and billing status in headers. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents and other MCP clients. Read the ScreenshotNeo API documentation for request options and response details.
curl -G "https://api.screenshotneo.com/v1/shot"
-d access_key=YOUR_API_KEY
--data-urlencode url=https://example.com/product
-o shot.webp
The free plan includes 1,000 screenshots per month with no card required; paid plans start at $5 for 3,000 shots. Visit ScreenshotNeo to learn more, or sign up free for 1,000 screenshots a month with no card.
Frequently Asked Questions
Can Puppeteer extract a value that is not visible on screen?
Yes. If the value exists in the rendered DOM or an attribute, extract it without requiring visibility; use visible: true only when visibility is part of the condition you need.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallShould I use page.content() to scrape the rendered value?
Use it to inspect or debug the rendered HTML. For routine extraction, a targeted $eval or $$eval avoids parsing the whole document.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

