Use Puppeteer when the data you need appears only after a page runs JavaScript or responds to browser interaction. Navigate to the page, wait for a signal tied to the data you need, extract only the required fields, validate the result, and close the browser. If the data is already present in the HTML or a direct JSON response, an HTTP request and parser are usually simpler.
When Puppeteer is the right tool
Puppeteer controls Chrome or Firefox through their supported automation interfaces, and runs headless by default. It is useful when browser execution, interaction, or rendered page state is necessary to reach the content. It is not a guarantee that a site permits scraping, nor does it make selectors or page behavior stable.
Before launching a browser, inspect whether a direct HTTP request already returns the HTML or JSON you need. Parsing that response avoids browser setup when no client-side rendering or interaction is required. Use Puppeteer when the browser changes the page in a way your task depends on.
Install Puppeteer and its browser
The normal puppeteer install path installs a compatible browser. puppeteer-core is the library-only alternative; choose it when you manage the browser separately. If your package manager or deployment environment blocks dependency install scripts, verify the browser installation explicitly using the current Puppeteer installation guide.
Recommended Free Tools
#1 Best Overall
npm install puppeteer
Run the scraper in an environment where the installed browser can launch. A package being present does not by itself confirm that its browser binary was installed or is available at runtime.
Build a scraper around page-specific readiness
The following illustrative example waits for product cards, extracts titles and links, rejects an empty result, and closes the browser even if navigation or extraction fails. Replace the URL and selectors after inspecting the target page; the selectors below are examples, not universal site structure.
import puppeteer from 'puppeteer';
const url = 'https://example.com/catalog';
const cardSelector = '.product-card';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (response && !response.ok()) {
throw new Error(`Unexpected page status: ${response.status()}`);
}
// Wait for the target page's actual content, not an arbitrary delay.
await page.locator(cardSelector).wait();
const records = await page.$$eval(cardSelector, cards =>
cards.map(card => ({
title: card.querySelector('.title')?.textContent?.trim() ?? '',
url: card.querySelector('a')?.href ?? ''
}))
);
if (records.length === 0) {
throw new Error('No product cards found');
}
console.log(records);
} finally {
await browser.close();
}
Use the current Puppeteer API documentation for version-specific details. The locator and extraction APIs are documented there; an arbitrary site’s selectors still need to be checked against its current DOM.
Choose a wait that proves the state you need
A page can be loaded as a document while its useful records are still being fetched or rendered. Choose a condition that corresponds to the task, then verify the resulting DOM or URL. A fixed sleep only says time passed; it does not establish that the needed state arrived.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →- Element appears: Use a locator wait or
waitForSelector(selector, { visible: true })when a particular result element is the readiness signal. - Custom DOM condition: Use
waitForFunctionwhen readiness means a minimum result count or a status field changing. - Document or URL changes: Use
waitForNavigationfor navigation, including History API URL changes in single-page applications. Start waiting before the action that triggers it. - Specific server response: Use
waitForResponsewith a narrow URL, method, or status predicate, then independently check that the expected UI state appeared. A request being sent does not prove the server accepted it. - Iframe appears: Wait for the frame and query inside it rather than searching the parent document for its contents.
- Network becomes quiet:
waitForNetworkIdlecan be appropriate for work such as capturing a page after late resources load, but network quiet does not prove the target data is correct.
Handle clicks that navigate
Register the navigation wait before clicking so the event cannot occur before Puppeteer begins waiting. Then inspect the response when one exists and verify the resulting URL or page content.
const [response] = await Promise.all([
page.waitForNavigation({ waitUntil: 'domcontentloaded' }),
page.locator('a.next-page').click()
]);
if (response && !response.ok()) {
throw new Error(`Unexpected status: ${response.status()}`);
}
if (!page.url().includes('/page/2')) {
throw new Error(`Unexpected URL after pagination: ${page.url()}`);
}
A same-document transition can produce a null response. Do not treat that alone as failure; check the expected URL or DOM state for the flow instead.
Select and extract only the data you need
CSS selectors work throughout Puppeteer’s selector APIs. Puppeteer also supports custom selector syntax for XPath, text, accessibility attributes, and Shadow DOM. Locators are the recommended normal interaction layer; they handle action preconditions and retry behavior. For direct querying after the page is ready, $, $$, $eval, and $$eval can locate elements or map DOM nodes to values.
- Prefer selectors anchored to meaningful attributes or content over brittle positional selectors when the page offers them.
- Keep extraction in the page context with
$evalor$$evalwhen practical, and return plain data rather than retaining browser element handles. - If you acquire an
ElementHandlewithwaitForSelector, dispose of it when finished. After a document replacement, query the new document instead of reusing handles from the old one. - Revalidate selectors when a site changes. The API can locate elements, but it cannot guarantee an arbitrary site’s DOM structure will remain stable.
Validate results and keep collection bounded
A script can finish without an exception and still return an error page, an empty list, or malformed fields. Make assumptions explicit and fail visibly rather than silently saving bad data.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Check the HTTP response status where available, and verify the resulting URL when redirects or navigation matter.
- Check that the expected records exist and that extracted fields match their expected formats.
- Distinguish an empty result from a timeout, unexpected page, or selector change in logs.
- Make pagination finite with a known stopping condition or a maximum page count; do not let a changed next-page control create an unbounded loop.
- Close the browser in a
finallyblock, and dispose lower-level handles that are no longer needed.
Use request interception only when needed
Request interception can provide control over network requests, but enabling it changes request handling: every intercepted request must be continued, responded to, aborted, or served by cache. If a request is left unresolved, it can stall the page and prevent the readiness condition from occurring. Use interception only when it serves a specific need and ensure every request path is handled.
Troubleshoot common failures
- Browser fails to launch: Check that installation scripts ran and that a compatible browser is installed in the runtime environment. If you use
puppeteer-core, account for the separately managed browser. - Wait times out: Confirm the selector or condition against the current page, whether the content is inside an iframe or shadow root, and whether the page reached the expected URL. Prefer the event or DOM state relevant to the task over increasing a blind delay.
- Navigation wait hangs or races: Start
waitForNavigationbefore the click or action. If the app updates without a document navigation, wait for the URL or DOM change instead. - Response wait matches unrelated traffic: Narrow the predicate by URL and, where useful, method or status. Then verify the expected page content independently.
- Response is non-OK or unexpected: Treat status, final URL, and expected content as separate checks; a request being made is not proof it succeeded.
- Extraction returns no records: Recheck the page state and selectors, and distinguish an empty result from a changed layout or an error page.
- Requests stop completing after interception is enabled: Ensure every intercepted request is continued, answered, aborted, or served from cache.
Respect access rules and legal boundaries
Review the site’s terms and access conditions before collecting data, and consider the data type, purpose, privacy and intellectual-property rules, and applicable jurisdiction. Obtain qualified legal advice for a consequential project. The legal position depends on the specific facts; browser automation does not grant permission.
The Internet Engineering Task Force’s September 2022 RFC 9309, Robots Exclusion Protocol, describes robots.txt rules as instructions requested of crawlers and states: “These rules are not a form of access authorization.” Robots.txt is therefore not a legal permission slip or a substitute for terms, authorization, or legal review.
The U.S. Supreme Court’s June 3, 2021 opinion in Van Buren v. United States interpreted “exceeds authorized access” under the U.S. Computer Fraud and Abuse Act in terms of obtaining information from computer areas off limits to the user. The case concerned a law-enforcement database; it did not decide that scraping any public website is lawful. See the Supreme Court opinion.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Or skip the browser setup
If your goal is a screenshot rather than structured data extraction, ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. This does not replace Puppeteer when you need to collect and transform page data.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options. Before the capture, it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo free.
Frequently asked questions
Does Puppeteer run headless?
Yes. Puppeteer’s official documentation says it runs headless by default; see the Puppeteer documentation for current browser and launch details.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Does robots.txt tell me whether scraping is legally allowed?
No. RFC 9309 says robots.txt rules are not a form of access authorization. Consider the site’s terms and applicable law as well.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




