Use Playwright to scrape an SPA by waiting for the application state you need, not merely for the browser’s first load event. Navigate with page.goto(), wait for a meaningful locator or page-state condition, then extract through locators (or browser-context evaluation when necessary). A single-page application can continue fetching data and replacing DOM nodes after domcontentloaded or load, so those milestones are only the beginning of a reliable workflow.
What makes single-page applications different?
A traditional page often delivers the content in the initial HTML response. An SPA may return a shell, load JavaScript bundles, fetch JSON, render components, and update the URL without another document navigation. If your scraper reads immediately after goto(), it can see an empty table, a loading label, or only the shell.
Playwright’s navigation options describe document lifecycle milestones. domcontentloaded means the document has been parsed; load waits for the load event. Neither proves that an SPA’s asynchronous request and rendering work is complete. The condition you wait for should be tied to the data you intend to extract.
A robust scraping sequence
- Launch an isolated browser context. Keep cookies, permissions and viewport settings explicit.
- Navigate to the route. Choose
domcontentloadedorloadas an initial milestone based on the site. - Wait for application evidence. Use a locator becoming visible, a loading status changing, a result count reaching a useful value, or another observable state.
- Interact through locators. Locator actions are retryable and auto-wait for actionable elements, which is safer than querying a stale element handle.
- Extract after readiness. Use locator text and attributes for ordinary fields; use locator evaluation or
page.evaluate()for browser-side processing. - Validate the result. Check that required fields are present and that the page did not show an error, access challenge or empty state.
Complete JavaScript example
The following script waits for a real result state before collecting cards. Replace selectors and the URL with those from the target application.
#1 Best Overall
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
const context = await browser.newContext({
viewport: { width: 1440, height: 900 }
});
const page = await context.newPage();
try {
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 30_000
});
const results = page.locator('[data-testid="result-card"]');
await page.locator('[data-testid="results-ready"]').waitFor({
state: 'visible',
timeout: 30_000
});
const count = await results.count();
if (count === 0) {
throw new Error('The application reported ready but returned no result cards');
}
const rows = await results.evaluateAll(cards => cards.map(card => ({
title: card.querySelector('[data-testid="title"]')?.textContent?.trim() ?? null,
price: card.querySelector('[data-testid="price"]')?.textContent?.trim() ?? null,
href: card.querySelector('a')?.href ?? null
})));
console.log(JSON.stringify(rows, null, 2));
} finally {
await browser.close();
}
})();
Run it with node scrape.js after installing a Playwright package and its browser. The results-ready marker is an example: use a selector or state that the application actually exposes.
Choosing the right readiness signal
Document milestones: domcontentloaded and load
Use domcontentloaded when you only need the parsed shell before starting an interaction. Use load when resources tied to the load event matter. Follow either with an application-specific wait whenever the desired content is rendered asynchronously.
Why networkidle is not a universal answer
Playwright documents networkidle as discouraged for general readiness: it considers an operation finished after no network connections for at least 500 ms. An SPA can keep a websocket, analytics request, polling loop or prefetch active even after the required content is visible; conversely, a quiet network can occur while the UI still shows a loading state. Treat network-idle as a narrow, site-specific signal rather than a blanket completion rule.
Locator and page-state conditions
Prefer conditions that describe the extraction target: a result heading visible, a spinner hidden, an error banner absent, a status changing to “Loaded,” or a known number of rows present. Locator APIs retry against current page state, so they remain useful when a framework re-renders the component.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →URL-aware synchronization
When clicking a link or submitting a form changes the route, wait for the expected URL and then wait for the destination content:
await Promise.all([
page.waitForURL('**/products/**'),
page.getByRole('link', { name: 'Products' }).click()
]);
await page.getByRole('heading', { name: 'Products' }).waitFor();
A URL transition alone does not establish that the SPA has finished rendering.
Rank #2
Extracting data from a changing DOM
Use locators for ordinary fields
Build locators from stable roles, labels, test IDs or semantic attributes. Read text with innerText() or textContent(), and attributes with getAttribute(). These calls resolve against the current DOM instead of depending on an element reference captured before a re-render.
Do not assume locator.all() waits
locator.all() returns the elements present immediately; it does not wait for a dynamic list to finish populating. First establish a useful stable condition, then collect the list:
const list = page.locator('[data-testid="result-card"]');
await expect(page.locator('[data-testid="loading"]')).toBeHidden();
await expect(list.first()).toBeVisible();
const cards = await list.all();
for (const card of cards) {
console.log(await card.innerText());
}
If the list can grow after the first batch, wait for a count or application status that defines the boundary you need, such as “showing 50 of 50.”
Use browser-context evaluation deliberately
page.evaluate() executes in the page’s browser context, separate from your Playwright script. Browser globals such as document exist there, and returned promises are awaited. Values from the Node.js context must be passed as serializable arguments:
const selector = '[data-testid="result-card"]';
const values = await page.evaluate((sel) => {
return [...document.querySelectorAll(sel)].map(node => ({
text: node.textContent?.trim() ?? '',
link: node.querySelector('a')?.href ?? null
}));
}, selector);
Keep complex business logic in your script where possible; use evaluation for DOM operations that are clearer or faster in the page.
Handling common SPA patterns
Loading indicators and empty states
Wait for the spinner to become hidden and for either a result or an intentional empty-state message. A hidden spinner alone can also mean an error, so inspect the error region before accepting an empty result.
Pagination and “load more” controls
Click the control through a locator, wait for the result count to increase, and stop when the control is disabled or absent:
const items = page.locator('[data-testid="item"]');
for (;;) {
const before = await items.count();
const more = page.getByRole('button', { name: /load more/i });
if (await more.count() === 0 || !(await more.isEnabled())) break;
await more.click();
await page.waitForFunction(
previous => document.querySelectorAll('[data-testid="item"]').length > previous,
before
);
}
Infinite scroll
Scroll only as far as needed, then wait for the list count or a “no more results” marker to change. Add a maximum-page or maximum-item limit so a faulty endpoint cannot create an unbounded job.
Client-side filtering and sorting
After changing a filter, wait for the visible result state to update rather than sleeping for an arbitrary duration. If the app exposes a result count or status text, assert its new value before extraction.
Route changes without full navigation
Some routers replace the URL while keeping the document. Pair waitForURL() with a locator for the new view, and do not rely on a second goto() unless you intentionally want a fresh document.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsRank #4
Timeouts, retries and reliability
Set a navigation timeout and a condition timeout that reflect the target’s normal behavior. A timeout should fail with context, not silently produce partial data. On failure, capture the URL, visible status text, a screenshot and (where permitted) the relevant HTML for diagnosis. Retry transient navigation or server failures with bounded exponential backoff; do not blindly repeat a deterministic selector failure.
Use a fresh context for independent identities, preserve a context when a login session is required, and close the browser in a finally block. Respect robots directives, terms, authentication boundaries and rate limits for the site you access. Playwright’s API documentation does not establish that any particular target permits automated extraction.
Troubleshooting empty or incorrect results
The script returns an empty array
- The selector may target a pre-render shell. Inspect the live DOM and wait for the application’s result marker.
- The list may still be changing. Wait for a count, status or hidden spinner before collecting it.
- The content may be inside an iframe or shadow root. Locate the correct frame or component boundary.
Timeout while waiting for a locator
- Confirm the route and selector in headed mode.
- Check for an error, consent dialog or authentication redirect blocking the view.
- Increase the timeout only after verifying that the condition is correct; a longer timeout cannot fix a wrong selector.
Data appears only after a click
Perform the click with a locator, wait for the resulting URL when it changes, and then wait for the newly revealed content. If the click opens a menu or dialog, assert that container is visible before querying inside it.
Works locally but fails in a job runner
Compare browser versions, viewport, timezone, locale, credentials and network access. Log the final URL and page errors. Use deterministic waits and avoid assumptions about headed rendering or local cache.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Duplicate or stale records
Frameworks may reuse nodes while updating data. Extract after the application reports the intended state, include a stable record ID where available, and deduplicate by that ID rather than by display text alone.
Performance and operating costs
Browser automation is heavier than an HTTP client because it executes JavaScript and maintains a full page. Reuse a browser process when jobs are trusted, but isolate contexts when cookies or permissions must not leak. Limit concurrency to what the target and your machine can support. Block unnecessary resources only when doing so does not remove data required for rendering. Cache results with a documented freshness window, and record the URL, timestamp, readiness condition and item count with each extraction.
Measure useful throughput as complete, validated records per minute—not pages started per minute. A fast script that captures loading shells needs reprocessing and costs more overall.
Or skip the browser setup
ScreenshotNeo provides a single-call website screenshot API and MCP server when you need a rendered visual rather than DOM records. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.
Use the API documentation at https://screenshotneo.com/docs/ for all options. cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Features include full-page and element capture, device presets, custom waits, CSS and JavaScript, request blocking, cookies and headers, geolocation, PDF controls, signed links, asynchronous webhooks, bulk capture and a usage API. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Start with a free ScreenshotNeo account.
When Playwright remains the right tool
Choose Playwright when the deliverable is structured data, when you must click through application state, or when readiness depends on content that only becomes available after interaction. Choose a screenshot API when the deliverable is a rendered image or PDF and you do not need to inspect every DOM record. In either case, define success as an observable, validated page state rather than a timer.
Frequently Asked Questions
Can Playwright scrape an SPA without waiting for JavaScript?
It can load the document, but reliable extraction requires waiting for the client-rendered state that contains the data. Document events alone are not sufficient for most asynchronous SPAs.
Recommended Free Tools
What should I log when a scraping run fails?
Log the final URL, timeout type, readiness condition, visible error text, item count and (where allowed) a diagnostic screenshot or HTML snapshot.
How do I avoid scraping an endlessly updating feed?
Define a stopping rule such as a target count, a page limit, a timestamp cutoff or an explicit end-of-results marker, and enforce it in the loop.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




