The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Playwright lets you drive Chromium, Firefox, and WebKit with the same browser, context, page, locator, and event APIs. A reliable scraper normally follows this sequence: launch a browser, create an isolated context, open a page, wait for a condition that proves the page is ready, locate data with user-facing locators, validate the result, and close everything in a finally block. The examples below use the standalone Playwright library in JavaScript rather than Playwright Test fixtures. Check the examples against the Playwright version installed in your project; the official documentation does not expose one stable version number for every page.
Only collect information you are allowed to access. Playwright does not grant permission to scrape a site, bypass authentication, defeat a CAPTCHA, or ignore a site’s terms and applicable law.
Install Playwright and run a first page
Create a project, install the library, and install at least one browser engine:
mkdir pw-scraper
cd pw-scraper
npm init -y
npm install playwright
npx playwright install chromium
This standalone script navigates to a benign example page, reads its title, and saves a screenshot. The browser is always closed, even when navigation or extraction fails.
Recommended Free Tools
#1 Best Overall
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
try {
const context = await browser.newContext();
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
console.log(await page.title());
await page.screenshot({ path: 'example.png', fullPage: true });
await context.close();
} finally {
await browser.close();
}
})();
The Page API documents the browser-to-context-to-page workflow and the basic page.screenshot() call: Playwright Page API.
Choose locators that survive page changes
Locators are the central piece of Playwright’s auto-waiting and retry-ability, according to the official locator guide. Prefer a locator that describes what a user sees or what the site explicitly promises, rather than a long CSS chain tied to today’s DOM.
Role and accessible-name locators
const heading = page.getByRole('heading', { name: 'Latest articles' });
await heading.waitFor();
const articleCards = page.getByRole('article');
const texts = await articleCards.evaluateAll(cards =>
cards.map(card => card.textContent?.trim() ?? '')
);
console.log(texts);
Useful built-in choices include getByRole, getByText, getByLabel, getByPlaceholder, getByAltText, getByTitle, and getByTestId. For buttons, links, headings, and form controls, a role plus accessible name is usually clearer than a class name.
Filter a repeated card before acting
When every product card has a similar button, first narrow the parent locator, then find the child control:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →const product = page.getByRole('listitem').filter({ hasText: 'Coffee grinder' });
await product.getByRole('button', { name: 'Add to cart' }).click();
This expresses the intended relationship and avoids clicking the first matching button elsewhere on the page. CSS and XPath remain available when a semantic locator or explicit test contract is unsuitable, but selectors such as div:nth-child(3) > span.item are coupled to structure that can change.
Extract only the fields you need
const rows = page.getByRole('row');
const records = await rows.evaluateAll(items => items.map(row => {
const cells = [...row.querySelectorAll('th,td')]
.map(cell => cell.textContent?.trim() ?? '');
return { name: cells[0] ?? '', value: cells[1] ?? '' };
}));
for (const record of records) {
if (!record.name || !record.value) {
throw new Error(`Invalid row: ${JSON.stringify(record)}`);
}
}
console.log(JSON.stringify(records, null, 2));
evaluateAll() runs a DOM operation over the elements currently matched by a locator. Keep the mapping focused, normalize whitespace, and validate required fields after extraction. The output depends on the target page’s markup; no selector is universal.
Rank #2
Wait for the page’s real readiness condition
Navigation completion is not the same as data readiness. A single-page application may render its list after the initial document loads. Wait for a meaningful condition instead of adding an arbitrary sleep.
Wait for a heading or list
await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
await page.getByRole('heading', { name: 'Catalog' }).waitFor();
const items = page.getByRole('listitem');
await items.first().waitFor();
const names = await items.allTextContents();
Wait for a selector or application state
await page.waitForSelector('[data-ready="true"]');
await page.waitForFunction(() => window.catalogLoaded === true);
Use a condition tied to the page’s contract: a heading, a row, a “loaded” marker, or a known application state. If the page’s list changes while you collect it, do not call locator.all() immediately and assume it waits. The Locator API documentation warns that all() returns current matches without waiting, so changing lists can produce unpredictable results. Wait for the list’s readiness condition first, then read it.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Navigation timeout and failure handling
page.setDefaultNavigationTimeout(45_000);
page.setDefaultTimeout(15_000);
try {
await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
} catch (error) {
console.error(`Navigation failed for ${targetUrl}:`, error.message);
throw error;
}
A timeout does not prove that a site is down: it can mean slow resources, a redirect loop, a consent wall, authentication, or a bot check. Record the URL and error, then decide whether to retry under the site’s permitted access rules.
Scrape a paginated list without losing your place
For a conventional “Next” link, collect one page at a time and stop when the control is disabled or absent. Scope each extraction to the current page and set a maximum page count as a safety guard.
const results = [];
const maxPages = 20;
for (let pageNumber = 1; pageNumber <= maxPages; pageNumber++) {
await page.getByRole('article').first().waitFor();
const pageRecords = await page.getByRole('article').evaluateAll(cards =>
cards.map(card => ({
title: card.querySelector('h2,h3')?.textContent?.trim() ?? '',
text: card.textContent?.trim() ?? ''
}))
);
results.push(...pageRecords);
const next = page.getByRole('link', { name: /next/i });
if (await next.count() === 0 || await next.isDisabled().catch(() => false)) break;
await Promise.all([
page.waitForLoadState('domcontentloaded'),
next.click()
]);
}
console.log(`Collected ${results.length} records`);
Some sites use a “Load more” button or infinite scrolling instead. Click the control and wait for the count to increase, or scroll only as far as the page’s documented behavior requires. Deduplicate records by a stable URL or identifier before writing output.
Keep users and sessions isolated with BrowserContexts
A BrowserContext is an isolated, incognito-like profile. Cookies, local storage, permissions, and other session state are separated, and contexts are designed to be fast and inexpensive to create. The browser-context documentation also shows how separate contexts model multiple users in one workflow.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
const browser = await chromium.launch();
try {
const guest = await browser.newContext();
const member = await browser.newContext();
const guestPage = await guest.newPage();
const memberPage = await member.newPage();
await guestPage.goto('https://example.com');
await memberPage.goto('https://example.com/account');
// Cookies or local storage created in one context are not shared with the other.
await guest.close();
await member.close();
} finally {
await browser.close();
}
Use one context when a workflow intentionally shares a login. Create separate contexts when testing guest/member behavior, processing independent accounts, or preventing one job’s cookies from affecting another. Isolation is a way to organize permitted sessions; it is not a way to bypass access controls.
Capture full-page, element, and in-memory screenshots
Save a full-page image
await page.screenshot({ path: 'catalog.png', fullPage: true });
Capture one element
const card = page.getByRole('article').first();
await card.screenshot({ path: 'first-card.png' });
Keep the image in memory
const buffer = await page.screenshot({ type: 'png' });
require('node:fs').writeFileSync('catalog-buffer.png', buffer);
The stable Page API covers these basic workflows. Playwright’s next-version screenshots guide is forward-looking; verify any option described there against the installed release before depending on it. A full-page image can be tall and expensive to process, while an element shot is smaller but requires a stable locator and a rendered element.
Wait for downloads and save them before closing the context
Start waiting for the download before clicking. The page emits its download event when the download starts; then save the completed file with saveAs().
const downloadPromise = page.waitForEvent('download');
await page.getByText('Download file').click();
const download = await downloadPromise;
const fileName = download.suggestedFilename();
await download.saveAs(`/absolute/path/output/${fileName}`);
Files associated with a browser context are deleted when that context closes, so save the file before closing the context. In production, validate the suggested filename and resolve it beneath an intended output directory rather than allowing path separators supplied by a remote page. See the Download API. The next downloads page is also forward-looking; prefer the stable API page unless you have verified your version.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsCombine navigation, extraction, and download in one script
This complete example visits a page, extracts visible article data, takes a screenshot, and saves a download if the page exposes the expected control.
const { chromium } = require('playwright');
const fs = require('node:fs/promises');
async function run(targetUrl) {
const browser = await chromium.launch();
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
try {
const page = await context.newPage();
await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.getByRole('article').first().waitFor({ timeout: 15_000 });
const articles = await page.getByRole('article').evaluateAll(nodes =>
nodes.map(node => ({
title: node.querySelector('h2,h3')?.textContent?.trim() ?? '',
text: node.textContent?.replace(/s+/g, ' ').trim() ?? ''
}))
);
if (articles.some(article => !article.title)) throw new Error('A record has no title');
await fs.writeFile('articles.json', JSON.stringify(articles, null, 2));
await page.screenshot({ path: 'page.png', fullPage: true });
const downloadLink = page.getByRole('link', { name: /download/i });
if (await downloadLink.count()) {
const pending = page.waitForEvent('download');
await downloadLink.click();
const download = await pending;
const safeName = download.suggestedFilename().replace(/[^a-zA-Z0-9._-]/g, '_');
await download.saveAs(`downloads/${safeName}`);
}
} finally {
await context.close();
await browser.close();
}
}
run('https://example.com').catch(error => {
console.error(error);
process.exitCode = 1;
});
Common failures and practical fixes
“Locator resolved to multiple elements”
Your locator is ambiguous. Add an accessible name, use filter({ hasText: ... }), or select a deliberately scoped parent before the child control. Do not blindly add first() unless the first match is genuinely the required one.
Rank #4
“Timeout exceeded” while waiting
Check the URL, redirects, authentication state, consent UI, and whether the expected role or accessible name exists. Inspect the page with a headed browser or save its HTML for diagnosis. Replace a guessed sleep with a condition that represents readiness, and increase the timeout only when the slower operation is expected.
An empty or partial collection
The list may still be rendering, may be virtualized, or may be inside an iframe. Wait for a specific row or application marker before reading it. If it changes while being read, avoid immediate all(); collect after the list stabilizes and validate the record count and required fields.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Clicks do not start the download
Start waitForEvent('download') before the click, ensure the locator targets the actual download control, and check whether a new tab or popup is opened instead. Save the download before closing its context.
Data differs between runs
Record the URL, timestamp, viewport, locale, and relevant response or page errors. Dynamic content, personalization, rate limits, and changing markup can all affect output. Keep concurrency within the target’s permitted limits and retry only errors that are safe to retry.
Performance, reliability, and operating choices
- Reuse a browser process carefully: launching one browser and creating contexts for independent jobs avoids accidental cookie sharing while keeping lifecycle management explicit.
- Choose the smallest result: extract fields rather than entire HTML documents, and capture an element instead of a full page when that is all you need.
- Make readiness observable: wait for a role, selector, count, or application state that proves the data exists; do not treat a fixed delay as a guarantee.
- Validate at the boundary: reject missing identifiers, malformed URLs, unexpected empty pages, and duplicate records before they reach downstream systems.
- Preserve diagnostics: keep structured logs for navigation failures and selected screenshots or HTML snapshots when permitted. Do not store credentials or personal data unnecessarily.
- Use bounded work: cap pages, records, retries, and output size so a changed site cannot create an unbounded job.
The reviewed Playwright documentation does not provide a universal speed or success-rate benchmark. The right browser, context, locator, and waiting strategy depends on the target page and your permitted workload.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
For a server-side screenshot rather than a local browser workflow, ScreenshotNeo accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Its cleanup steps accept the cookie or consent banner as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools.
See the ScreenshotNeo API documentation for authentication and options. This is a one-call WebP example:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The equivalent Python and Node.js calls are:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector, delay or network idle, request and resource blocking, headers, cookies, user agent, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with the parameter names used by other screenshot APIs. Every feature is on every plan: Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try the 1,000 monthly screenshots without a card.
Frequently asked questions
Should I use Playwright Test or the standalone library?
Use the standalone library when your program is a scraper, data pipeline, or one-off automation job. Use Playwright Test when you also want its test runner, fixtures, assertions, and reporting. The examples here intentionally use the library API.
Can Playwright scrape content rendered inside an iframe?
Yes, when you can legally access it: obtain a frame locator or frame reference and use locators within that frame. The page locator and the iframe’s document are separate scopes.
Why does a screenshot show a consent banner?
Playwright captures what its page rendered. Locate and handle a permitted consent flow before capture, or hide a known selector only when doing so accurately represents your intended result. A screenshot service such as ScreenshotNeo can perform its documented consent and widget cleanup before capture.
How should I schedule large scraping jobs?
Queue bounded jobs, isolate sessions with contexts, limit concurrency for the target, persist checkpoints, and make retries idempotent. Add monitoring for empty results and markup changes instead of assuming every run succeeded.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

