October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk8 min

How to Scrape Data from React, Vue, and Angular Websites

Learn how to find data behind React, Vue, and Angular pages, decide between direct requests and browser rendering, and validate your results responsibly.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If a scraper returns empty HTML from a React, Vue, or Angular site, first find where the data is delivered: in the initial response, embedded in a script, or in a later network request. Use a plain HTTP request when it exposes the required data; use a headless browser when the page must execute JavaScript or reach a browser-specific state before the data appears. The framework name alone does not determine the right method.

Why a JavaScript website can look empty to a scraper

An HTTP-only scraper receives the server’s response but does not run the page’s JavaScript. A site may return the content in its initial HTML, include it as structured data in a script, or send a minimal app shell that JavaScript later fills with content. React, Vue, and Angular sites can use any of these patterns; server-side rendering and pre-rendering can also put content in the initial response. Google describes the app-shell distinction and notes that not all bots execute JavaScript (Google Search Central).

The browser’s live DOM is not the same artifact as the original response source. If text appears in the browser but not in your HTTP response, it may have been inserted later or fetched separately.

Diagnose where the target data comes from

  1. Fetch the URL without rendering. Save the response body and search it for the target text. Inspect script elements for embedded JSON or other structured data. Compare the downloader’s response with the page source from an ordinary HTTP client, rather than assuming the live browser DOM was in the original HTML. Scrapy recommends this comparison when diagnosing missing content (Scrapy: Dynamic content).
  2. Inspect the browser’s network activity. Open the browser developer tools, select Network, reload the page, and inspect requests whose responses contain the target records. Look for JSON or other text responses from fetch/XHR requests, and check whether the data is instead in the original response or a JavaScript resource.
  3. Choose the least complex permitted source. If the content is in HTML, use an HTML parser. If it is embedded JSON, parse that representation. If a relevant JSON request supplies the records, reproduce that request and parse its response where doing so is appropriate and permitted. These approaches avoid coordinating a browser when the browser is not needed; they do not guarantee that a site’s endpoint will remain stable or that you are authorized to use it.
  4. Render only when needed. If the useful content depends on script execution, interactions, or browser state and reconstructing its request is impractical, automate a real browser and inspect its rendered DOM.
  5. Wait for the data, not just the page load. Wait for a selector tied to the content or another observable condition, then validate the result. A fixed delay can help diagnose timing, but elapsed time alone does not establish that the page is ready.
  6. Validate representative records. Check that expected fields exist, inspect item counts, and detect empty or error states. Revisit your selectors and request assumptions if routes, lazy-loaded content, or page updates change.

Choose an extraction method

What you observe Start with Why
Target text is present in the raw response HTML HTTP client and HTML selectors The response already contains the data; JavaScript execution is unnecessary for that extraction.
Target data is embedded in a script Extract and parse the embedded representation Scrapy documents extracting JavaScript text and parsing JSON-like content where practical.
A network request returns the target data as JSON or another structured response Reproduce the relevant request and parse its response This can return the data directly without rendering the full page.
Data appears only after scripts run or browser-specific state is reached Playwright or another headless browser A browser exposes the rendered DOM and can support necessary page interactions.
You need crawl orchestration across many pages with occasional browser rendering Scrapy with a browser integration Scrapy supports browser-rendering integrations for cases where ordinary downloading is insufficient.

The practical trade-offs are implementation complexity, completeness, runtime and resource needs, and sensitivity to changes in the site’s behavior. There is no universal speed or success-rate figure that applies to all sites and extraction methods.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Playwright when the rendered page is the source

Playwright’s Page API provides locator-based waits and browser interaction (Playwright Page API). The example below waits for a site-specific result selector, then reads its rendered text. Replace the URL and selector with values confirmed in the target page’s DOM; this example extracts one element’s text, not a complete crawl.

import { chromium } from 'playwright';

const url = 'https://example.com/products';
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'domcontentloaded' });

  const results = page.locator('[data-testid="product-list"]');
  await results.waitFor({ state: 'visible', timeout: 15000 });

  const text = await results.innerText();
  if (!text.trim()) {
    throw new Error('Product list appeared but contained no text');
  }
  console.log(text);
} finally {
  await browser.close();
}

Install Playwright and its browser binaries according to the official Playwright getting started guide. The selector in the sample is illustrative, not a universal convention: inspect the site’s actual rendered DOM and choose a selector that identifies the data you need.

Wait for a meaningful condition

domcontentloaded indicates that the initial document has been parsed; it does not guarantee that client-side data fetching or rendering has finished. The example therefore waits for the result container. If the container appears before its records are added, wait for a more specific child or validate the record count. Playwright documents selector and locator waits in its Page API.

For a one-off diagnostic, a short fixed delay can help reveal whether content arrives later, but it is brittle: slow responses can outlast it, while fast pages make it waste time. Prefer a selector, expected record condition, or another signal tied to the content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle pagination and lazy loading deliberately

A visible first set of records may not represent the full dataset. Determine whether the page uses pagination, an infinite-scroll trigger, or a separate request for more records. For each additional page or batch, wait for the relevant state change and validate that new records appeared. Do not treat a successful page load as proof that all records were collected.

Extract and validate without hiding failures

  • Parse HTML with selectors when the data is in HTML; parse JSON as JSON rather than treating it as markup.
  • Check required fields on representative records and reject results that are empty or structurally different from what you expect.
  • Record enough context to diagnose failures, such as the requested URL, response status, whether the target selector appeared, and the resulting record count.
  • Recheck selectors and request assumptions when the site changes its layout, route behavior, or data-loading pattern.
  • Keep requests within the target’s access rules and avoid treating a missing result as a reason to bypass a login, CAPTCHA, or other access control.

Respect robots.txt and access conditions

Check the site’s robots.txt, terms, access controls, and applicable legal requirements before collecting data. RFC 9309, the IETF Robots Exclusion Protocol standard published in September 2022, says: “These rules are not a form of access authorization.” Robots.txt communicates crawler access preferences; an allowed path is not permission to access protected content. See RFC 9309. Legal outcomes depend on the facts and jurisdiction, particularly for authenticated, personal, copyrighted, or otherwise restricted material.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It can capture rendered pages as PNG, JPEG, WebP, or PDF; it is useful when the job is to capture a page or PDF, rather than extract structured records from a site. A screenshot is an image of the page, not a replacement for a JSON response or parsed DOM when your task requires structured data.

One GET request returns a screenshot. This cURL example uses the documentation’s URL and saves a WebP response; add your API key and choose an output format appropriate to your request:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for parameters and options. Before capture, it can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and other MCP clients.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; all features are available on every plan. Sign up for ScreenshotNeo’s free plan.

Common problems and fixes

Symptom Likely cause What to check or change
HTTP response has no target text, but the browser does The content is added after the initial response or fetched separately. Inspect Network responses for the data request. Reproduce a suitable structured request, or render the page if browser execution is required.
Browser automation returns an empty result The script may still be loading data, the selector may not match, or the page may be in an empty/error state. Inspect the rendered DOM and network activity; wait for a content-specific selector and validate fields and record count.
A wait times out The selector may be wrong, the content may not load, or the page may require another state or interaction. Confirm the selector in the live DOM, inspect failed requests and page errors, and verify whether the content is actually available under your current access.
Only some records are collected Pagination, lazy loading, or incremental data fetching may limit the initial result set. Identify the pagination or loading mechanism, process each batch, and check that new records appear before continuing.
The scraper breaks after a site update Selectors or request assumptions may depend on changeable site behavior. Reinspect the response and network flow, update the extraction logic, and keep validation checks that detect missing or malformed data.

Further reading

For a broader Python scraping reference, Web Scraping with Python, 3rd Edition by Ryan Mitchell was published by O’Reilly in February 2024. The publisher lists 352 pages and describes it as an intermediate-to-advanced book covering JavaScript scraping and crawling through APIs (O’Reilly listing). It is optional background, not a prerequisite for the workflow above.

Frequently Asked Questions

Do React, Vue, or Angular always require a headless browser to scrape?

No. Choose based on where the target data is delivered: the initial response, embedded data, a separate request, or only the rendered page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does an allowed path in robots.txt mean I have permission to scrape it?

No. RFC 9309 says robots.txt rules are not access authorization; check the site’s access conditions separately.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.