Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The fastest Playwright scraper is usually not the one with the most aggressive blocking or the highest page count in parallel. It is the one that waits only for data it needs, avoids safe-to-skip work, reuses browser resources deliberately, and measures completed records rather than raw request speed. Start by replacing broad waits with content-specific readiness checks, then test selective request blocking, explicit context lifecycles and carefully increased concurrency against the same target pages.
1. Measure a trustworthy baseline first
Before changing code, run the current scraper against the same URLs, browser version, machine and extraction rules. Record total elapsed time, navigation time, time waiting for dynamic content, parsing time, memory use, failed pages and the number of complete records. A faster run that silently misses lazy-loaded fields is not an optimization.
- Keep a fixed URL sample that represents normal, slow and JavaScript-heavy pages.
- Log timestamps before navigation, after navigation, after the content-ready condition and after extraction.
- Record HTTP failures, timeouts, CAPTCHA or bot-check pages and empty results separately.
- Compare cold visits and repeat visits; caching can change the result.
Playwright’s documentation does not publish a universal scraper speedup, concurrency limit or benchmark. Treat every change below as an experiment and keep it only when required data remains correct.
2. Wait for the data, not for the whole internet
page.goto() uses load by default. The Page API also supports commit, domcontentloaded and networkidle; the latter means no network connections for at least 500 ms and is explicitly discouraged as a general readiness signal. Choose the earliest event that leaves your extraction data available, then wait for a locator or response that represents the actual data requirement.
#1 Best Overall
See the Playwright Page API for the current navigation contract.
Use a specific locator for server-rendered content
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('[data-product-card]').first().waitFor({ state: 'visible', timeout: 15_000 });
const products = await page.locator('[data-product-card]').evaluateAll(cards =>
cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim(),
price: card.querySelector('.price')?.textContent?.trim()
}))
);
If the page fills a table after an API call, wait for a row, a result count or a response associated with that call instead of sleeping for an arbitrary duration.
Use response conditions when the API result is the contract
await page.goto(url, { waitUntil: 'commit' });
await page.waitForResponse(response =>
response.url().includes('/api/search') && response.ok(),
{ timeout: 20_000 }
);
Do not stack a long fixed delay on top of navigation and a content wait unless the target demonstrably needs it. If content is variable, a bounded condition is both faster on quick pages and safer on slow ones.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors3. Remove requests selectively, and understand routing costs
Routing can continue, abort or fulfill requests. If your extractor never uses images, video or advertising calls, abort only those categories after verifying that the target does not use them for layout, lazy loading or application logic.
await context.route('**/*', async route => {
const type = route.request().resourceType();
if (type === 'image' || type === 'media' || type === 'font') {
await route.abort();
} else {
await route.continue();
}
});
This is not universally safe: CSS, fonts and scripts may affect selectors, pagination or lazy loading. Start with one resource class and validate extracted fields.
Two routing caveats
- Enabling routing disables HTTP cache. A route that saves transfers on the first visit can make repeat navigation slower, so test both cold and warm runs. The cache behavior is documented in the BrowserContext API.
- Context routing does not intercept requests handled by a service worker. If interception is essential, consult Playwright’s service-worker guidance and block service workers only when that matches the behavior you need to scrape.
Use the Network guide to inspect requests before deciding what can be removed.
4. Reuse the browser, isolate the sessions
For a batch, launch one browser process, create explicit contexts for independent sessions, create pages inside those contexts and close each resource in a predictable order. Playwright describes contexts as isolated, fast and cheap to create. browser.newPage() is a convenience API suited to short, single-page scenarios, not a substitute for lifecycle control in a production scraper. See the Browser API and browser-context isolation guide.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →import { chromium } from 'playwright';
const browser = await chromium.launch();
try {
for (const url of urls) {
const context = await browser.newContext();
const page = await context.newPage();
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('[data-record]').first().waitFor({ timeout: 10_000 });
// extract and persist the record here
} finally {
await context.close();
}
}
} finally {
await browser.close();
}
Reuse a context when cookies and local storage should persist for a session. Use separate contexts when credentials, locale or state must not leak between jobs.
5. Increase concurrency as a controlled experiment
Multiple isolated contexts can run in one browser, but Playwright does not specify a safe number for arbitrary sites. Increase parallelism gradually while watching throughput, memory, browser crashes, timeouts, HTTP errors and the target site’s behavior. The useful setting depends on page weight, your machine and the site’s limits.
import pLimit from 'p-limit';
const limit = pLimit(4); // starting point for an experiment, not a universal limit
const results = await Promise.all(urls.map(url => limit(() => scrapeOne(browser, url))));
Raise the limit only if completed, correct records per minute improve without a disproportionate rise in failures. Reduce it when pages compete for memory or the site begins returning challenges. Playwright’s fixtures documentation covers isolated contexts and efficient browser use; it does not prescribe a concurrency number: Fixtures API.
Rank #3
6. Separate site latency from your own code
Instrument navigation, waits, browser evaluation, parsing and persistence independently. A slow third-party response is different from a slow selector loop or synchronous file write. The best-practices guide discusses controlling third-party responses in tests; for scraping, use that idea to identify local overhead, not as a reason to mock the real data source you need.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Reduce local overhead
- Extract arrays in one
evaluateAllcall instead of making one round trip per element. - Stream results to storage in batches rather than serializing a large in-memory object after every page.
- Reuse compiled parsing logic and avoid repeated full-page HTML transfers when locators provide the needed fields.
- Set bounded timeouts and classify failures so one bad page does not stall the queue.
7. A complete, measured scraper pattern
import { chromium } from 'playwright';
async function scrapeOne(browser, url) {
const context = await browser.newContext();
const page = await context.newPage();
const started = Date.now();
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
const cards = page.locator('[data-product-card]');
await cards.first().waitFor({ state: 'attached', timeout: 15_000 });
const data = await cards.evaluateAll(nodes => nodes.map(node => ({
name: node.querySelector('.name')?.textContent?.trim() ?? null,
price: node.querySelector('.price')?.textContent?.trim() ?? null
})));
return { url, data, elapsedMs: Date.now() - started };
} finally {
await context.close();
}
}
const browser = await chromium.launch();
try {
// Begin with a small limit, then compare correctness and throughput.
const queue = urls.map(url => scrapeOne(browser, url));
const results = await Promise.allSettled(queue);
console.log(results);
} finally {
await browser.close();
}
In production, replace the unbounded Promise.all with a queue or limiter, add retries only for transient errors, and persist each successful record before moving on.
8. Troubleshooting slow or incomplete runs
The scraper waits until timeout
The selector may be wrong, hidden, or never rendered for a particular URL. Capture the URL and page state, verify the locator in a headed run, and wait for the actual result condition rather than a generic network event.
Blocking assets breaks extraction
Remove the route, identify which resource supplies the missing data, then block only confirmed nonessential types. A script, stylesheet or service worker may be part of the page’s data path.
Repeat visits became slower after routing
Routing disables HTTP cache. Compare a no-route baseline and decide whether reduced transfer outweighs lost cache reuse for your workload.
Parallel pages cause failures or memory pressure
Lower concurrency, close contexts in finally blocks, and monitor browser and host memory. There is no documented universal limit.
Pages show a bot check or blank result
Classify the outcome instead of retrying indefinitely. Respect the target’s terms and access controls; a faster loop cannot make an unavailable page valid data.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
For jobs that need a rendered screenshot rather than DOM extraction, ScreenshotNeo provides a single website-screenshot API request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers identify the page verdict and billing status.
Use cURL (see the ScreenshotNeo API documentation):
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same call in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Every plan includes its features; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Is networkidle ever appropriate?
It can describe a specific page’s behavior, but Playwright discourages it as a general readiness test because background traffic can keep a page waiting. Prefer a condition tied to the fields you extract.
Best Value
Should I block images on every scrape?
No. Block only resources you have verified are unnecessary for that target and extraction path; images or scripts may trigger lazy loading or application behavior.
How many Playwright pages can run at once?
No universal safe number is documented. Increase concurrency gradually and choose the level that preserves completed-record correctness and acceptable resource use.
Frequently Asked Questions
Is networkidle ever appropriate?
It can describe a specific page’s behavior, but Playwright discourages it as a general readiness test. Prefer a condition tied to the fields you extract.
Should I block images on every scrape?
No. Block only resources verified to be unnecessary for that target and extraction path.
How many Playwright pages can run at once?
There is no universal safe number; increase concurrency gradually while monitoring correctness, failures and resource use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

