Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Yes—Playwright is a strong choice for scraping pages whose useful data appears only after JavaScript runs, a user interaction occurs, or a session is established. Use it as a controlled browser, not as a way around access controls: define permission and data limits first, choose direct HTTP when it is sufficient, isolate each job in its own browser context, wait for the data state rather than a timer, and scale with bounded concurrency, retries, checkpoints and observability.
What Playwright adds to a scraper
A conventional HTTP client downloads HTML. Playwright drives Chromium, Firefox or WebKit, so the page can execute JavaScript, issue client-side requests, set cookies and local storage, and expose the same rendered interface a user sees. That makes it useful for dashboards, search results, infinite lists and authenticated workflows that an HTML parser cannot interpret alone.
Browser rendering is more expensive and can introduce more failure modes than a direct request. Start with the lightest transport that can legally and reliably return the fields you need.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches| Situation | Preferred approach | Why |
|---|---|---|
| Stable public JSON or HTML response | HTTP client or the site’s official API | Lower latency and resource use, simpler retries |
| Data appears after JavaScript rendering | Playwright page plus resilient locators | Runs the application that creates the data |
| Authorized login, filters or button-driven state | Playwright browser context | Maintains cookies and browser state while you interact |
| A stable request contains the required data | Observe or call that response through the Network API | Avoids extracting from fragile presentation markup |
Playwright’s locators are the central abstraction for finding elements, waiting for actionability and retrying against a changing DOM. Prefer semantic locators such as roles, labels, text, placeholders, alt text, titles and configured test IDs over long CSS or XPath chains.
#1 Best Overall
Start with an ethical and operational scope
Whether a crawl is lawful depends on the target, your purpose and the applicable jurisdiction. Before opening a browser, document:
- Operator and purpose: identify the organization running the job and why each field is needed.
- Target and frequency: list exact domains, paths, schedule, expected volume and a stop condition.
- Permission boundaries: read the site’s terms, machine-readable directives and rate-limit guidance. Do not cross an authentication boundary without authorization, and prefer an official API or export when one exists.
- Data policy: collect only necessary fields, define retention and deletion dates, restrict access to raw pages and session artifacts, and redact personal data from logs.
- Change and denial handling: stop when the operator blocks access, consent requirements change, throttling persists or the returned schema no longer matches your contract.
Never design retries, proxies or concurrency to defeat a bot check or access control. A permission that covers one account, tenant or purpose does not automatically cover another.
Install a reproducible Playwright runtime
Pin the Playwright package and browser revision in your project and CI image. A fresh virtual environment prevents one job’s cookies or local storage from leaking into another.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallnpm init -y
npm install playwright
npx playwright install chromium
Set the target explicitly rather than embedding credentials or a private domain in source control:
export TARGET_URL='https://your-authorized.example/catalog'
export OUTPUT='items.json'
node scrape.mjs
A complete JavaScript scraper
The following example creates a new context for one job, waits for a meaningful heading and list, extracts cards with semantic locators, records a stable key, and retries only transient navigation failures. Replace the example selectors with labels or test IDs documented for your authorized target.
import { chromium } from 'playwright';
import fs from 'node:fs/promises';
const url = process.env.TARGET_URL;
if (!url) throw new Error('Set TARGET_URL');
const maxAttempts = 3;
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
async function run() {
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
locale: 'en-US',
timezoneId: 'UTC'
});
const page = await context.newPage();
page.setDefaultTimeout(15000);
try {
let response;
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
try {
response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 30000
});
break;
} catch (error) {
if (attempt === maxAttempts) throw error;
await sleep(500 * 2 ** (attempt - 1));
}
}
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
await page.getByRole('heading', { name: /catalog|products/i }).waitFor();
const cards = page.getByRole('article');
await cards.first().waitFor();
const count = await cards.count();
const items = [];
for (let i = 0; i < count; i++) {
const card = cards.nth(i);
const name = (await card.getByRole('heading').first().innerText()).trim();
const price = (await card.getByText(/$|€|£/).first().innerText()).trim();
const link = await card.getByRole('link').first().getAttribute('href');
items.push({ name, price, link });
}
// Validate the contract before writing a checkpoint.
for (const item of items) {
if (!item.name || !item.link) throw new Error('Schema validation failed');
}
await fs.writeFile(process.env.OUTPUT ?? 'items.json', JSON.stringify(items, null, 2));
console.log(JSON.stringify({ status: 'ok', count: items.length }));
} finally {
await context.close();
await browser.close();
}
}
run().catch(error => {
console.error(JSON.stringify({ status: 'error', message: error.message }));
process.exitCode = 1;
});
The example deliberately uses domcontentloaded only as a navigation milestone. It then waits for the heading and the first card, which are evidence that the data needed for extraction exists. Navigation readiness is not data readiness.
Synchronize with state, not arbitrary sleeps
Playwright auto-waits for actionability when it clicks or fills controls, and web-first assertions retry until their condition is true. Use those mechanisms for extraction as well:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →- Wait for a specific heading, table, card, empty-state message or application response.
- For a request-backed view, wait for the expected response while performing the action, then validate its status and shape.
- Use a short delay only for a documented debounce or animation that has no observable state; never make a long fixed sleep your primary synchronization strategy.
- Playwright documents
networkidleas discouraged for testing. Pages with analytics, sockets or polling may never become idle, so an assertion tied to the required data is safer.
Dynamic lists need special care: locator.all() does not wait for a list to stabilize. Wait for a known item, count threshold or loading indicator to disappear, then count and read items as in the example.
Build selectors that survive UI changes
Prefer semantic locators
Use getByRole, getByLabel, getByText, getByPlaceholder, getByAltText, getByTitle and a configured test ID. Scope a locator to a meaningful container before filtering it by stable text or an attribute.
Avoid presentation coupling
Generated class names, deeply nested CSS and absolute XPath expressions often change when a frontend build changes. If you control the target application, ask for stable test IDs or a documented data attribute. Keep selectors in one module so a UI change is repaired once.
Rank #3
Validate the extracted schema
Reject missing identifiers, malformed URLs and impossible values before writing a checkpoint. A successful browser run that silently emits empty fields is a data failure, not a success.
Pagination, infinite scroll and network responses
Numbered pagination
After extracting a page, save its URL or cursor and the last stable key. Click the next control only when it is enabled, wait for the old content to be replaced or for a response to complete, and stop when the control is absent or the cursor repeats. Deduplicate records by a stable identifier.
Infinite scroll
Record the item count, trigger the documented load-more action or scroll, then wait until the count increases or an explicit end marker appears. Stop after a maximum page count and retain the checkpoint so a crash does not restart from the beginning.
Network interception
If the required fields arrive in a stable, authorized response, observe that response or route it through the browser context rather than scraping rendered text. Collect only the fields needed, avoid logging authorization headers or unrelated payloads, and preserve the application’s expected behavior. Browser-context routing and request listeners are useful for this class of work.
Isolate sessions and credentials
A BrowserContext isolates cookies, local storage and session state. Create one per job, tenant or account boundary. Do not share authenticated state between unrelated customers. If you deliberately persist state, encrypt the storage file, limit its lifetime and permissions, and delete it when the job’s retention period ends. Supply credentials through a secret manager or environment injection, never through source code or screenshots.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scale without turning load into abuse
Bound concurrency
Each browser page consumes CPU and memory. Use a small, measured worker pool rather than launching unbounded tabs. Set per-navigation and per-step timeouts, and lower concurrency when latency, error rates or throttling rise.
Retry selectively
Retry transient DNS failures, connection resets and temporary server errors with capped exponential backoff and jitter. Do not retry permission failures, repeated access denials, consent changes or a stable schema error indefinitely. Every retry should be visible in metrics.
Cache and checkpoint
Cache responses or completed records for a declared time-to-live when the use case permits. Checkpoint after each page or cursor, include the input and software version, and make restarts idempotent. Caching reduces duplicate requests and helps respect rate limits.
Measure the workload
Track pages completed, records accepted, latency, retry count, HTTP status classes, duplicate rate, schema failures, throttling and browser resource use. Pin Playwright and browser versions so a future upgrade is a controlled change. For visual comparisons, keep operating-system and browser versions consistent.
Recommended Free Tools
Failure modes and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Timeout waiting for a locator | Wrong selector, consent overlay, slow data request or changed UI | Inspect the accessible role/name, handle an authorized consent flow, wait for the specific response or update the selector; do not just increase the timeout. |
| Empty list with a successful navigation | Data loads after navigation or requires a filter/session | Wait for the list’s meaningful state, verify context cookies and filters, and capture the response that should contain the records. |
locator.all() returns too few items |
The list was still rendering | Wait for a stable item, count threshold or end-of-loading marker, then enumerate. |
| HTTP 401 or 403 | Missing authorization or access is not permitted | Stop and obtain documented permission or use the official integration. Do not automate around the denial. |
| Repeated 429 or throttling | Concurrency or frequency exceeds the operator’s limit | Honor the stated limit, reduce workers, back off and resume from a checkpoint. |
| Records change between runs | Live data, rotating experiments or unpinned browser/runtime | Record timestamps and versions, pin the runtime, and define whether the job requires a snapshot or current values. |
| Browser crashes under load | Too many pages, heavy media or memory leaks | Lower concurrency, block unnecessary resource types when authorized, close pages promptly and monitor memory per worker. |
Performance, reliability and cost decisions
Measure your own authorized workload; official Playwright documentation does not provide a universal throughput or success percentage. The main cost drivers are browser startup, page complexity, JavaScript execution, media, concurrency and retries. Reusing a browser process while creating fresh contexts usually avoids repeated startup work without sacrificing session isolation. Keep a hard upper bound on pages, bytes and elapsed time for each job.
Best Value
Reliability is more than a green process exit. Require a non-empty result when one is expected, validate every record, persist checkpoints, classify failures and alert on schema drift. A clean stop on an access denial is preferable to an apparently complete but unauthorized crawl.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP or PDF. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
Use the documented parameters and options when a rendered image or PDF—not structured records—is the actual requirement. Features include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector/delay/network idle, request or resource blocking, headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Common parameter names from other screenshot APIs also work.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for response formats and options. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients, so an authorized AI agent can request captures without you maintaining browser automation.
There is a free allowance of 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to try it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

