What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
If a scraper returns empty HTML from a React, Vue, or Angular site, first find where the data is delivered: in the initial response, embedded in a script, or in a later network request. Use a plain HTTP request when it exposes the required data; use a headless browser when the page must execute JavaScript or reach a browser-specific state before the data appears. The framework name alone does not determine the right method.
Why a JavaScript website can look empty to a scraper
An HTTP-only scraper receives the server’s response but does not run the page’s JavaScript. A site may return the content in its initial HTML, include it as structured data in a script, or send a minimal app shell that JavaScript later fills with content. React, Vue, and Angular sites can use any of these patterns; server-side rendering and pre-rendering can also put content in the initial response. Google describes the app-shell distinction and notes that not all bots execute JavaScript (Google Search Central).
The browser’s live DOM is not the same artifact as the original response source. If text appears in the browser but not in your HTTP response, it may have been inserted later or fetched separately.
Diagnose where the target data comes from
- Fetch the URL without rendering. Save the response body and search it for the target text. Inspect script elements for embedded JSON or other structured data. Compare the downloader’s response with the page source from an ordinary HTTP client, rather than assuming the live browser DOM was in the original HTML. Scrapy recommends this comparison when diagnosing missing content (Scrapy: Dynamic content).
- Inspect the browser’s network activity. Open the browser developer tools, select Network, reload the page, and inspect requests whose responses contain the target records. Look for JSON or other text responses from fetch/XHR requests, and check whether the data is instead in the original response or a JavaScript resource.
- Choose the least complex permitted source. If the content is in HTML, use an HTML parser. If it is embedded JSON, parse that representation. If a relevant JSON request supplies the records, reproduce that request and parse its response where doing so is appropriate and permitted. These approaches avoid coordinating a browser when the browser is not needed; they do not guarantee that a site’s endpoint will remain stable or that you are authorized to use it.
- Render only when needed. If the useful content depends on script execution, interactions, or browser state and reconstructing its request is impractical, automate a real browser and inspect its rendered DOM.
- Wait for the data, not just the page load. Wait for a selector tied to the content or another observable condition, then validate the result. A fixed delay can help diagnose timing, but elapsed time alone does not establish that the page is ready.
- Validate representative records. Check that expected fields exist, inspect item counts, and detect empty or error states. Revisit your selectors and request assumptions if routes, lazy-loaded content, or page updates change.
Choose an extraction method
| What you observe | Start with | Why |
|---|---|---|
| Target text is present in the raw response HTML | HTTP client and HTML selectors | The response already contains the data; JavaScript execution is unnecessary for that extraction. |
| Target data is embedded in a script | Extract and parse the embedded representation | Scrapy documents extracting JavaScript text and parsing JSON-like content where practical. |
| A network request returns the target data as JSON or another structured response | Reproduce the relevant request and parse its response | This can return the data directly without rendering the full page. |
| Data appears only after scripts run or browser-specific state is reached | Playwright or another headless browser | A browser exposes the rendered DOM and can support necessary page interactions. |
| You need crawl orchestration across many pages with occasional browser rendering | Scrapy with a browser integration | Scrapy supports browser-rendering integrations for cases where ordinary downloading is insufficient. |
The practical trade-offs are implementation complexity, completeness, runtime and resource needs, and sensitivity to changes in the site’s behavior. There is no universal speed or success-rate figure that applies to all sites and extraction methods.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Use Playwright when the rendered page is the source
Playwright’s Page API provides locator-based waits and browser interaction (Playwright Page API). The example below waits for a site-specific result selector, then reads its rendered text. Replace the URL and selector with values confirmed in the target page’s DOM; this example extracts one element’s text, not a complete crawl.
import { chromium } from 'playwright';
const url = 'https://example.com/products';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'domcontentloaded' });
const results = page.locator('[data-testid="product-list"]');
await results.waitFor({ state: 'visible', timeout: 15000 });
const text = await results.innerText();
if (!text.trim()) {
throw new Error('Product list appeared but contained no text');
}
console.log(text);
} finally {
await browser.close();
}
Install Playwright and its browser binaries according to the official Playwright getting started guide. The selector in the sample is illustrative, not a universal convention: inspect the site’s actual rendered DOM and choose a selector that identifies the data you need.
Rank #2
Wait for a meaningful condition
domcontentloaded indicates that the initial document has been parsed; it does not guarantee that client-side data fetching or rendering has finished. The example therefore waits for the result container. If the container appears before its records are added, wait for a more specific child or validate the record count. Playwright documents selector and locator waits in its Page API.
For a one-off diagnostic, a short fixed delay can help reveal whether content arrives later, but it is brittle: slow responses can outlast it, while fast pages make it waste time. Prefer a selector, expected record condition, or another signal tied to the content.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Handle pagination and lazy loading deliberately
A visible first set of records may not represent the full dataset. Determine whether the page uses pagination, an infinite-scroll trigger, or a separate request for more records. For each additional page or batch, wait for the relevant state change and validate that new records appeared. Do not treat a successful page load as proof that all records were collected.
Extract and validate without hiding failures
- Parse HTML with selectors when the data is in HTML; parse JSON as JSON rather than treating it as markup.
- Check required fields on representative records and reject results that are empty or structurally different from what you expect.
- Record enough context to diagnose failures, such as the requested URL, response status, whether the target selector appeared, and the resulting record count.
- Recheck selectors and request assumptions when the site changes its layout, route behavior, or data-loading pattern.
- Keep requests within the target’s access rules and avoid treating a missing result as a reason to bypass a login, CAPTCHA, or other access control.
Respect robots.txt and access conditions
Check the site’s robots.txt, terms, access controls, and applicable legal requirements before collecting data. RFC 9309, the IETF Robots Exclusion Protocol standard published in September 2022, says: “These rules are not a form of access authorization.” Robots.txt communicates crawler access preferences; an allowed path is not permission to access protected content. See RFC 9309. Legal outcomes depend on the facts and jurisdiction, particularly for authenticated, personal, copyrighted, or otherwise restricted material.
Rank #4
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It can capture rendered pages as PNG, JPEG, WebP, or PDF; it is useful when the job is to capture a page or PDF, rather than extract structured records from a site. A screenshot is an image of the page, not a replacement for a JSON response or parsed DOM when your task requires structured data.
One GET request returns a screenshot. This cURL example uses the documentation’s URL and saves a WebP response; add your API key and choose an output format appropriate to your request:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteBest Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters and options. Before capture, it can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and other MCP clients.
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; all features are available on every plan. Sign up for ScreenshotNeo’s free plan.
Common problems and fixes
| Symptom | Likely cause | What to check or change |
|---|---|---|
| HTTP response has no target text, but the browser does | The content is added after the initial response or fetched separately. | Inspect Network responses for the data request. Reproduce a suitable structured request, or render the page if browser execution is required. |
| Browser automation returns an empty result | The script may still be loading data, the selector may not match, or the page may be in an empty/error state. | Inspect the rendered DOM and network activity; wait for a content-specific selector and validate fields and record count. |
| A wait times out | The selector may be wrong, the content may not load, or the page may require another state or interaction. | Confirm the selector in the live DOM, inspect failed requests and page errors, and verify whether the content is actually available under your current access. |
| Only some records are collected | Pagination, lazy loading, or incremental data fetching may limit the initial result set. | Identify the pagination or loading mechanism, process each batch, and check that new records appear before continuing. |
| The scraper breaks after a site update | Selectors or request assumptions may depend on changeable site behavior. | Reinspect the response and network flow, update the extraction logic, and keep validation checks that detect missing or malformed data. |
Further reading
For a broader Python scraping reference, Web Scraping with Python, 3rd Edition by Ryan Mitchell was published by O’Reilly in February 2024. The publisher lists 352 pages and describes it as an intermediate-to-advanced book covering JavaScript scraping and crawling through APIs (O’Reilly listing). It is optional background, not a prerequisite for the workflow above.
Frequently Asked Questions
Do React, Vue, or Angular always require a headless browser to scrape?
No. Choose based on where the target data is delivered: the initial response, embedded data, a separate request, or only the rendered page.
Recommended Free Tools
Does an allowed path in robots.txt mean I have permission to scrape it?
No. RFC 9309 says robots.txt rules are not access authorization; check the site’s access conditions separately.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




