Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Playwright when the data appears only after JavaScript runs. Launch a browser, open a fresh context, navigate with page.goto(), wait for a meaningful locator or the API response that supplies the page, then extract text, attributes, or structured JSON. Prefer role-, label-, text-, and test-ID-based locators over brittle CSS or XPath. Close the context and browser in a finally block so every run releases its resources.
The Playwright scraping workflow
A JavaScript-rendered page is an application, not just an HTML document. The initial response may contain an empty shell while scripts request products, comments, prices, or account data. Playwright runs a real browser, so your scraper can observe the rendered DOM or capture the network response that populated it.
- Install the Node.js package and browser binaries.
- Launch Chromium (or another installed browser), create an isolated
BrowserContext, and open aPage. - Navigate with
page.goto(). Navigation waits for the page’sloadevent by default. - Synchronize with a meaningful locator, an assertion, or a response promise instead of an arbitrary sleep.
- Extract visible data with locators, or parse a structured API response.
- Close the context and browser in cleanup code.
Install Playwright and run a first scraper
Install the library and browsers
npm init -y
npm install playwright
npx playwright install
Use npx playwright install chromium when the project only needs Chromium. Browser binaries are separate from the JavaScript package, so a package install alone is not sufficient on a new machine or CI runner.
Minimal rendered-DOM example
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
try {
await page.goto('https://example.com');
const heading = await page.getByRole('heading').first().textContent();
console.log({ heading: heading?.trim() });
} finally {
await context.close();
await browser.close();
}
Save this as an ES module (for example, set "type": "module" in package.json) and run it with Node.js. The finally block is important when navigation or extraction throws an error.
#1 Best Overall
Choose selectors that survive redesigns
Locators are Playwright’s central mechanism for auto-waiting and retryability. They describe what a user can recognize instead of tying the scraper to a particular nesting structure.
| Preferred locator | Typical use | Why it is resilient |
|---|---|---|
getByRole() |
Buttons, headings, links, rows | Uses the accessible role and name visible to users |
getByText() |
Distinct visible labels | Follows user-facing copy when it is stable |
getByLabel() |
Form controls | Connects to the control’s accessible label |
getByPlaceholder() |
Search and input fields | Uses an explicit input hint |
getByAltText() |
Images with meaningful alternative text | Uses the image’s accessibility contract |
getByTitle() |
Elements with a title attribute | Targets an intentional tooltip/title value |
getByTestId() |
Internal data contracts | Remains stable when visual markup changes, if the site maintains the test ID |
Use CSS or XPath only when a stable contract requires it, such as a documented data attribute or a selector supplied by the site owner. Avoid chains such as div:nth-child(3) > span; a harmless layout change can silently return the wrong record.
Extracting a list
const cards = page.getByRole('article');
const count = await cards.count();
const products = [];
for (let i = 0; i < count; i++) {
const card = cards.nth(i);
products.push({
name: (await card.getByRole('heading').textContent())?.trim(),
price: (await card.getByText(/$|€|£/).first().textContent())?.trim()
});
}
console.log(products);
When a locator matches multiple elements, use count() and nth(), or use a locator assertion that expresses the expected state. Do not assume the first match is the desired record unless that is part of the page’s contract.
Free tools Windows power users keep installed
One-click scans. No signup required.
Wait for dynamic content without guesswork
Wait for a meaningful locator
After navigation, wait for the element that proves the data you need is present. Locator actions and assertions automatically wait for actionability and retry transient states.
await page.goto('https://example.com/products');
const rows = page.getByRole('row');
await rows.nth(1).waitFor({ state: 'visible' });
const firstName = await rows.nth(1).getByRole('cell').first().textContent();
A locator wait is more informative than a fixed delay: it ends as soon as the required state exists and times out when the page cannot produce it.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Synchronize with the action that reveals data
await page.goto('https://example.com/products');
const loadButton = page.getByRole('button', { name: 'Load products' });
await loadButton.click();
await page.getByRole('row').nth(1).waitFor({ state: 'visible' });
const visibleNames = await page.getByRole('row').allTextContents();
Do not use a generic fixed sleep as your primary synchronization method. Generic networkidle waiting and page.waitForSelector() are discouraged in Playwright’s testing guidance because background requests can make them ambiguous; an application-specific locator or response is clearer for a scraper too.
Capture the API response behind the page
If the interface is backed by JSON, response capture is often cleaner than parsing formatted text. Create the response promise before the click or navigation that triggers the request, then await it after the action.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
try {
await page.goto('https://example.com/products');
const responsePromise = page.waitForResponse('**/api/products');
await page.getByRole('button', { name: 'Load products' }).click();
const response = await responsePromise;
if (!response.ok()) {
throw new Error(`Products request failed: ${response.status()}`);
}
const data = await response.json();
console.log(data);
} finally {
await context.close();
await browser.close();
}
Use a predicate when several requests share a path:
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/products') &&
response.request().method() === 'GET' &&
response.status() === 200
);
Attach listeners when you need an audit trail of all traffic:
page.on('request', request => {
if (request.url().includes('/api/')) console.log('request', request.method(), request.url());
});
page.on('response', response => {
if (response.url().includes('/api/')) console.log('response', response.status(), response.url());
});
Response capture gives you the server’s structured payload, but it does not automatically grant permission to use an undocumented endpoint. Preserve required authentication and respect the site’s access rules.
Rank #3
Control requests, headers, and resources
Routing lets you inspect, modify, fulfill, or abort matching requests. Every intercepted request must be continued, fulfilled, or aborted; leaving it unresolved stalls the page.
Skip images for a text-only run
await context.route('**/*', async route => {
const type = route.request().resourceType();
if (type === 'image' || type === 'font' || type === 'media') {
await route.abort();
} else {
await route.continue();
}
});
Install the route before navigation. Blocking resources can reduce bandwidth and speed extraction, but do not block a resource needed to create the data or to satisfy a selector.
Inspect or alter a request
await context.route('**/api/products', async route => {
const request = route.request();
console.log(request.method(), request.url(), request.headers());
await route.continue({
headers: { ...request.headers(), 'x-scrape-run': 'catalog' }
});
});
Use fulfill() for a controlled fixture or abort() for unwanted traffic. Keep such behavior limited to targets you are authorized to automate.
Keep sessions isolated with BrowserContexts
A non-persistent context is an independent session: cookies, permissions, and storage do not leak into another context, and non-persistent contexts do not write browsing data to disk. Create one context per account, locale, or concurrent job when those sessions must remain separate.
const publicContext = await browser.newContext({ locale: 'en-US' });
const accountContext = await browser.newContext({ locale: 'en-US' });
const publicPage = await publicContext.newPage();
const accountPage = await accountContext.newPage();
try {
// Run independent workflows here.
} finally {
await publicContext.close();
await accountContext.close();
}
Use a persistent profile only when the workflow explicitly requires durable login state. For repeatable scraping, fresh contexts make cookies and permissions predictable.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Observe WebSocket-backed data
Some dashboards never fetch a conventional JSON response; they receive updates over WebSockets. Listen for the socket and inspect sent or received frames.
page.on('websocket', socket => {
console.log('WebSocket opened:', socket.url());
socket.on('framesent', frame => console.log('sent', frame));
socket.on('framereceived', frame => console.log('received', frame));
socket.on('close', () => console.log('WebSocket closed'));
});
Frame formats are application-specific. Record only the fields you need, and expect reconnects or heartbeat messages in long-running jobs.
Pagination, retries, and data quality
Paginate by the site’s own controls
Prefer a “Next” button or documented cursor over guessing page numbers. After each click, wait for a response or for the first row to change, then extract the page. Keep a set of seen IDs or URLs to prevent an accidental loop.
Validate before writing
- Require a stable key such as an ID, canonical URL, or SKU.
- Normalize whitespace, but keep the original URL and raw response when auditability matters.
- Check that a response has the expected content type and fields before calling
json(). - Write checkpoints after each page so a timeout does not discard earlier results.
Retry narrowly
Retry transient navigation or network failures with a bounded attempt count and increasing delay. Do not blindly retry a permission failure, a persistent HTTP error, or a selector timeout; log the URL, status, and failed condition so the target can be investigated.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
Executable doesn't exist |
Browser binaries were not installed | Run npx playwright install (or install the browser you launch). |
Empty text after goto() |
Data arrives after the initial load | Wait for a meaningful locator or capture the response triggered by the page. |
| Timeout waiting for a locator | Wrong role/name, consent overlay, changed markup, or failed request | Inspect the accessible name, observe requests and responses, and verify the page state before changing the selector. |
| Response promise times out | The promise was created after the click, the URL pattern is wrong, or the request is made only once during navigation | Create the promise before the triggering action and log matching requests to confirm the pattern. |
| Scraper returns duplicate or mixed-account data | Pages share cookies or storage | Use separate BrowserContexts and close each one after its job. |
| Page hangs after routing | A route handler did not continue, fulfill, or abort | Ensure every branch completes one of those operations. |
| API returns unauthorized | Authentication, headers, or cookies are required | Use an authorized session, supply permitted headers/cookies, and do not attempt to bypass access controls. |
Performance and reliability decisions
- Use one browser with multiple contexts for parallel independent sessions; launching a browser for every URL adds startup overhead.
- Reuse a page for a sequence when cookies and navigation state should persist; use a new context when isolation matters more than reuse.
- Block unnecessary resources only after confirming they are not part of the data path.
- Prefer API responses when the page exposes the exact structured records you need; use DOM extraction when the user-visible representation is the source of truth.
- Log evidence: target URL, navigation status, selector or response condition, elapsed time, and failure details.
- Bound every wait and close contexts in cleanup code so one broken page cannot consume all workers.
Scraping permissions and responsible use
Playwright documents browser automation mechanics, not whether a particular target may be scraped. Before running a job, review the site’s robots.txt, terms of service, authentication requirements, rate limits, copyright and privacy obligations, and the laws that apply to your location and the target. A technically successful request can still be unauthorized. Keep request rates reasonable, collect the minimum necessary data, protect credentials, and provide a way to stop a job when the site changes its policy.
Best Value
Or skip the browser setup
If you need a rendered image or PDF rather than extracted records, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and can return PNG, JPEG, WebP, or PDF without you managing Playwright browser binaries.
One request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for all parameters. The same endpoint works from Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Or Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
Other controls include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/orientation/page ranges, HTML/CSS rendering, custom JavaScript and CSS, pre-capture clicks, hidden selectors, waits for a selector/delay/network idle, request and resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, image resizing, chosen cache TTL, signed image links, asynchronous jobs with signed webhooks, up to 100 URLs per bulk call, a usage API, an OpenAPI specification, and compatible parameter names used by other screenshot APIs.
Every feature is on every plan: 1,000 shots per month are free with no card; paid plans are $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, or $249 for 1,000,000. Yearly billing gives two months free. Sign up for the free 1,000-shot plan when an API or MCP workflow is a better fit than maintaining a browser.
Frequently Asked Questions
Can Playwright extract data that is rendered inside an iframe?
Yes. Identify the frame and create locators within that frame; keep the frame selection explicit because a page can contain several frames with similar content.
How can I make a scraper reproducible when the site changes?
Store the selector or response condition used, retain a small raw sample, and fail loudly when required fields disappear instead of silently writing partial records.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallShould I parse the DOM or call the page’s API directly?
Use the DOM when the visible, user-facing result is authoritative. Capture the API response when it contains the same records in a stable structured form and your use is authorized.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

