Free tools Windows power users keep installed
One-click scans. No signup required.
Use Puppeteer when the content you need appears only after a browser runs JavaScript or responds to an interaction. It controls Chrome or Firefox so your JavaScript program can navigate a page, wait for the right state, interact with elements, and extract the resulting content. For static pages, a direct HTTP request may be simpler; Puppeteer does not grant permission to collect a site’s data.
What Puppeteer does—and when to use it
The Puppeteer project describes it as “a JavaScript library which provides a high-level API to control Chrome or Firefox over the DevTools Protocol or WebDriver BiDi.” It runs headless by default. That makes it useful for pages whose content or behavior depends on browser-side JavaScript, or where you must perform an action before the content is available. It is a browser automation library you can use to build a scraper, not a dedicated scraping appliance.
As an Amazon Associate I earn from qualifying purchases.
If the information is already available in the initial HTML or through a documented endpoint you are allowed to use, a browser may be unnecessary overhead. Choose the least complex permitted method that returns the data you need.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteFor target-site access rules and applicable requirements, check the specific site’s published policies and the rules relevant to your work. Keep collection limited to what you need. Puppeteer itself does not authorize access or bypass restrictions.
#1 Best Overall
Choose and install the right package
Choose between the two packages based on who manages the browser installation:
| Package | Browser setup | Best fit | Operational note |
|---|---|---|---|
puppeteer |
Downloads a compatible Chrome during installation. | You want the package-managed browser setup. | If your package manager blocks install scripts, the browser may not be downloaded. |
puppeteer-core |
Does not download Chrome with the library. | You manage or configure the browser separately. | You must provide a browser installation and configure Puppeteer to use it. |
Install the full package for the simplest initial setup:
npm install puppeteer
Or install the library without its managed browser download:
npm install puppeteer-core
If Chrome is missing because an install script was blocked, allow the relevant install script under your package manager’s policy or use Puppeteer’s documented manual browser-install route:
npx puppeteer browsers install
Browser availability and package-manager configuration can differ between environments. Do not assume a browser downloaded on a developer machine is also present in a clean CI or deployment environment.
A complete scraper: navigate, wait, extract, close
This Node.js example uses the package-managed browser, waits for a page element before extracting text, verifies the result, and closes the browser even if navigation or extraction fails. Replace the example URL and selector with a site and element you are permitted to access.
const puppeteer = require('puppeteer');
async function main() {
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
const response = await page.goto('https://example.com', {
waitUntil: 'domcontentloaded',
});
if (!response) {
throw new Error('Navigation did not return a response');
}
if (!response.ok()) {
throw new Error(`Page returned HTTP ${response.status()}`);
}
const heading = page.locator('h1');
await heading.wait();
const text = await heading.map(element => element.textContent);
const value = text.trim();
if (!value) {
throw new Error('The expected heading was present but contained no text');
}
console.log(value);
} finally {
await browser.close();
}
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
The sequence is: launch a browser, create a page, navigate to a URL with its scheme (such as https://), wait for the element relevant to your task, read its text, validate the value, and close the browser. A successful navigation call alone does not prove that the expected content loaded.
Find and interact with the right content
Prefer locators for actions
Puppeteer’s page-interactions guide recommends locators for interacting with page elements. Locators automatically wait for the element to be present and for the state needed to perform an action. For example:
const button = page.locator('button[type="submit"]');
await button.click();
Use a selector that describes the element on the target page, then check the resulting value or page state. A selector copied from an unrelated site’s markup is not evidence that the target uses the same structure.
Understand selector options
CSS selectors work by default. Puppeteer’s custom selector syntax also supports text, accessibility attributes, XPath, and Shadow DOM access. If a standard CSS query cannot reach the content, check whether it is inside a shadow root or a frame before concluding that the page has no matching element.
Read text after the relevant state is reached
Use a locator or an appropriate wait before extraction. Reading immediately after navigation can return an empty value if client-side rendering has not produced the target element yet. Validate both that the selector matched and that its extracted text or other value is usable.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Wait for the condition your task needs
Different pages need different readiness conditions. Puppeteer’s Page API provides navigation, selector, response, and network-idle waits. Choose the signal tied to the content or action you need rather than relying on an arbitrary fixed delay.
- Element appears: wait for the target selector, then extract or act on it.
- Element becomes visible: wait for visibility when hidden elements should not count as ready.
- A particular response arrives: wait for the relevant response when the data depends on a request.
- Navigation completes: wait for navigation when a click or other action takes the page elsewhere.
- Network becomes idle: use a network-idle condition only when it is a meaningful signal for the page. Ongoing background requests can make it a poor fit.
The default selector wait timeout is 30 seconds unless changed. A timeout means the expected condition was not met in that interval; it does not by itself tell you whether the selector is wrong, the page is still loading, or the content is elsewhere.
Avoid the click/navigation race
If a click triggers navigation, register the navigation wait at the same time as the click. Waiting only after clicking risks missing the navigation event:
await Promise.all([
page.waitForNavigation(),
page.locator('a.next-page').click(),
]);
Check the response and extracted data
page.goto() returns a response when navigation produces one; inspect its status when the page’s HTTP result matters. A response status does not establish that the expected content is present, so also wait for the relevant element and validate the extracted value. Handle the case where navigation returns no response instead of treating it as a successful data capture.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11For production scripts, keep browser closure in a finally block. Otherwise, a thrown navigation, timeout, or extraction error can skip cleanup and leave the browser process running.
Capture a screenshot or create a PDF
Screenshots can help inspect what the browser actually rendered when a selector fails or a page looks different from its source HTML. Puppeteer also supports PDF generation for an HTML page:
await page.screenshot({ path: 'page.png', fullPage: true });
await page.pdf({ path: 'page.pdf' });
page.pdf() uses print CSS by default, so the PDF may not look like the page’s screen layout. It creates a PDF from the current HTML page; that is different from downloading or parsing an existing PDF document. The headless shell cannot navigate directly to a PDF document.
Troubleshoot common failures
- Browser executable is missing: The package installation may have run with install scripts blocked. Allow the appropriate script or run
npx puppeteer browsers install, then verify the browser is available in the same environment as your script. - Extraction is empty: The page may not yet have reached the state that exposes the content, the selector may not match, or the content may live in a frame or Shadow DOM. Wait for the expected element and inspect the page structure.
- A selector wait times out: Confirm the selector against the actual page, check whether the target becomes visible only after an interaction, and determine whether the element is in a frame or shadow root. Increase a timeout only when the content legitimately needs longer; a longer timeout will not fix a selector that can never match.
- A click succeeds but navigation is missed: Register
page.waitForNavigation()alongside the click withPromise.all, rather than beginning the wait afterward. - The response is an error: Inspect the navigation response and its status. Do not treat a loaded error page as valid scraped content; validate the expected selector and extracted value as well.
- The script hangs or leaves browser processes behind after an error: Put browser closure in
finallyso cleanup runs on both success and failure. - PDF output differs from the visible screen: Remember that
page.pdf()renders with print media CSS by default. The headless shell also cannot navigate directly to an existing PDF document.
Or skip the browser setup
If your task is to capture a website rather than build a custom browser interaction flow, ScreenshotNeo offers a screenshot API and MCP server. One GET request returns a screenshot or PDF; this cURL example saves a WebP screenshot:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
Frequently Asked Questions
Can Puppeteer scrape a page that requires a login?
Puppeteer can automate browser interactions, but whether you may access or collect content behind a login depends on the site’s rules and applicable requirements. The documentation cited here does not establish permission for any particular site.
Does Puppeteer run only in headless mode?
No. It runs headless by default; browser launch configuration can be used when a visible browser is needed.
Can Puppeteer scrape every website?
No universal claim is justified. Sites differ in rendering, access rules, and page structure, and Puppeteer does not guarantee that a particular page is accessible or that its content can be collected.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




