Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
When a scraper gets an almost-empty page from a React, Vue, or Angular site, the app may not have rendered its data yet. First check whether the fields you need are already in an accessible API response or embedded page data; if not, use a browser such as Playwright and wait for the actual content—not just a navigation event—to appear. The framework name alone does not tell you which approach a particular URL needs.
Why a scraper may get an empty page
A basic HTTP request downloads the document but does not run its client-side JavaScript. A single-page app (SPA) may initially return a small HTML shell and script references; the browser then runs those scripts, fetches data, handles the route, and updates the DOM. A scraper that reads only the initial response can therefore miss content visible in a regular browser.
Check the actual URL rather than assuming that every React, Vue, or Angular page works the same way. These frameworks are clues that client rendering may be involved, not proof that a page is rendered only in the browser. Compare the initial document response with the DOM after the page appears. If the needed text or records are in the original response, a browser may be unnecessary; if they arrive later or require browser state or interaction, rendering is likely needed.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteInspect the page before choosing a method
- Open the page in a regular browser. Note the exact URL, route, and content you need to collect.
- Compare source and rendered DOM. Check the initial document response or page source, then inspect the live DOM after the content appears. Search for the fields you need, including serialized or hydration data embedded in the document.
- Inspect Network activity. In the browser developer tools, filter for fetch/XHR requests. Open likely responses and see whether they contain the needed fields, rather than only unrelated configuration or partial data.
- Determine whether state or interaction matters. If the request requires a client-side route, a session, a click, or other browser state, record those steps before choosing an extraction strategy.
If a useful response or embedded payload is available and can appropriately be requested, extracting it directly may be simpler than parsing rendered markup. That approach still depends on discovering and maintaining the relevant request or payload. Check the target site’s terms and access rules; the technical availability of an endpoint does not establish permission to use it.
#1 Best Overall
Choose between direct data, a browser, and a hybrid
| Approach | Best fit | Trade-off |
|---|---|---|
| Direct response or embedded data | The fields are present in an accessible API response or page payload, and no browser-only state is needed. | You must identify the relevant request or payload and handle changes to it. |
| Browser-rendered DOM | The data depends on JavaScript execution, client-side routing, browser state, or interaction. | You must run a browser and manage readiness, runtime, and failures. |
| Hybrid | A browser is needed to reach a state, but the data is then available in requests made by the page. | There are more moving parts, and you must verify the request flow and permitted use. |
Decide based on where the fields appear, whether authentication or interaction is required, operational complexity, and how sensitive the extraction is to page or request changes. There is no neutral speed, cost, or success-rate benchmark established for these approaches here, so choose based on the target rather than a universal performance ranking.
Render the app with Playwright when needed
Playwright supports Chromium, Firefox, and WebKit. Its official Browser documentation recommends creating a browser context and then a page explicitly in production code and test frameworks. The one-step browser.newPage() convenience is intended for short, single-page scenarios. A context gives you a clearer place to manage page lifetime and browser state. See the Playwright Browser API and Page API.
Here is a runnable Node.js example for a public page. It waits for a target-specific selector, extracts matching text, checks for an empty result, and closes the browser even if navigation or extraction fails. Replace the URL and selector with ones confirmed in the target page’s rendered DOM.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
try {
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 30000,
});
// Replace this with a selector for the actual data you need.
const results = page.locator('[data-testid="product-name"]');
await results.first().waitFor({ state: 'visible', timeout: 15000 });
const names = await results.allTextContents();
const cleanNames = names.map(name => name.trim()).filter(Boolean);
if (cleanNames.length === 0) {
throw new Error('The page loaded but no product names were found.');
}
console.log(JSON.stringify({ url: page.url(), names: cleanNames }, null, 2));
} finally {
await context.close();
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Install Playwright and its browser binary in your project environment using the official browser installation instructions. Playwright releases are paired with browser binaries; after upgrading Playwright, reinstall the required browser if the installed binary is missing or incompatible. In a container or CI job, include browser installation in the environment setup rather than assuming a local browser is already present.
Wait for the data, not a generic lifecycle event
The example uses domcontentloaded to begin after the initial document is parsed, then waits for the expected content. That selector is illustrative: choose an observable condition tied to the data you intend to extract.
- Wait for a selector when a stable element appears only after the target content renders.
- Wait for expected text when a specific label or record signals that the right view is ready.
- Wait for a known response when a specific data request is the clearest signal. Playwright’s Page API supports observing page events and requests.
A route change does not prove that its data has arrived. Likewise, load and network-idle signals are not universal readiness checks: an app may update after navigation, while polling or other long-lived requests may prevent the network from becoming idle. Browserless’s technical guide discusses these pitfalls for SPAs: Scraping React, Vue & Angular SPAs.
Rank #3
Give the readiness condition a finite timeout and treat failure as a useful result to diagnose, not permission to scrape an empty page. Record enough context—such as the URL, failure stage, and whether the expected selector appeared—to distinguish a changed layout from a slow or failed load.
Extract and validate records
Prefer stable, meaningful selectors or the underlying data response when appropriate; the framework name does not reveal the target’s DOM structure. Before accepting a run, validate that results are nonempty and include the fields your downstream task needs. If the page should contain multiple records, check a reasonable expected count or range for that particular page. Save the URL and retrieval time with the output so you can investigate later changes. These checks are implementation guidance, not results of a test against a particular site.
For a direct-response approach, inspect the request in the browser’s Network panel, identify the response fields you need, and reproduce the request only where permitted. Treat undocumented request formats as potentially changeable. For a hybrid, use the browser to reach the required route or state, then inspect the page’s requests to see whether a response supplies the records more directly than the DOM.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Troubleshoot common failures
- Only the app shell appears: The initial document may not contain the rendered data. Compare it with the live DOM and inspect fetch/XHR responses or embedded data; use browser rendering if the content requires JavaScript.
- The script times out waiting for a selector: The selector may be wrong, the route may not have loaded, or the target content may have changed. Confirm the selector in the live DOM and inspect the current URL and page state before increasing a timeout.
- The route changed but results are empty: Navigation finished before the app’s data update. Wait for an element, text, or response tied to the target records, then validate that extraction returned data.
- Network idle never arrives: Ongoing background requests may keep the page active. Replace generic idle waiting with a target-specific condition.
- Playwright cannot launch a browser: Check that the browser binary for the installed Playwright release exists in the current environment. Follow the official browser installation guidance after package upgrades or when preparing a container.
- The page works manually but not in automation: The scraper may be missing a required route, session, or interaction. Compare the steps and requests made in the regular browser with those made by the automated page; do not assume a framework-specific fix will apply to every site.
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than extract structured records, ScreenshotNeo is a screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. It is not a substitute for extracting and validating structured fields from a page.
For a screenshot, the cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options. Cookie banners, popups, and chat widgets are removed before capture; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up to get 1,000 free screenshots a month with no card.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Keep access, reliability, and cost in view
There is no single browser setup or readiness condition that works for every SPA. Direct extraction can avoid browser runtime when the required data is already available; browser rendering is useful when the page must execute scripts or establish state. A browser path adds infrastructure and the need to manage timeouts, browser binaries, and page lifetimes. Track empty results and failed readiness checks instead of silently treating them as successful captures. Before automating access, review the target site’s terms and applicable access rules.
For another perspective on SPA extraction, SparkProxy’s guide describes inspecting responses and choosing a method based on what the page exposes: How to Scrape Single Page Applications (SPAs). Prerender.io addresses a different problem: its integration documentation explains rendering and caching crawler-facing versions of a publisher’s own SPA, rather than collecting data from other sites: Prerender.io’s SPA integration documentation.
Best Value
Frequently Asked Questions
Does using React, Vue, or Angular mean I always need a browser scraper?
No. Check the particular page: the data may already be in its initial response or an embedded payload, even when the site uses one of those frameworks.
Can I use Firefox or WebKit instead of Chromium with Playwright?
Yes. Playwright documents Chromium, Firefox, and WebKit as browser choices; install the binary paired with the Playwright release you use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

