Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a real browser to run the SPA, wait for the field’s populated state, and then extract either the API response that supplied it or the rendered DOM. A normal HTTP request often returns only an application shell. The most reliable workflow is to discover the request carrying the custom field, capture it before the interaction that triggers it, and parse its JSON. If the value exists only after client-side formatting or an interaction, scope a locator to the correct record and read the rendered element.

Why a normal HTTP request misses custom fields

React, Vue and Angular applications commonly deliver a small HTML shell plus JavaScript bundles. The browser then requests records, applies authentication and state, renders components, and reveals fields after a tab click, search, scroll or “load more” action. An HTTP client such as requests or curl sees the shell unless you separately reproduce those data calls.

There are two useful extraction layers:

  • Network/API layer: capture the JSON response containing the record and custom field. This is usually less brittle than selectors tied to presentation markup.
  • Rendered-DOM layer: let the browser finish rendering, then read text, attributes or links from a stable element. Use this when the value is computed in the browser or is not present in a usable response.

Check the site’s robots directives, terms, authentication requirements, privacy and copyright obligations, rate limits and applicable law before collecting data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Map the page state before writing a scraper

  1. Identify the route and record key. Record the URL pattern, record ID and the container that represents one record.
  2. Find the field’s label and state. Note whether the custom field is visible immediately or appears after a Details tab, search, pagination, “load more” control or scrolling.
  3. Observe the browser’s requests. In developer tools, filter Fetch/XHR calls while performing the interaction. Look for JSON containing the record ID or field name.
  4. Record the trigger and completion signal. The trigger may be navigation or a click; completion should be a known response or a locator becoming visible, not an arbitrary delay.

Keep authentication cookies, local storage and headers in the same browser context that performs navigation. Otherwise the page can look logged out even when your interactive browser is authenticated.

Preferred method: capture the SPA’s JSON response

Register the response listener before navigation or the click that causes the request. The following Playwright script waits for a GET to /api/records, parses the payload, and preserves a null value when a field is absent.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();

const responsePromise = page.waitForResponse(response =>
  response.url().includes('/api/records') &&
  response.request().method() === 'GET' &&
  response.status() === 200
);

await page.goto('https://example.com/records', { waitUntil: 'domcontentloaded' });
const response = await responsePromise;
const payload = await response.json();

for (const record of payload.records ?? []) {
  console.log({
    id: record.id,
    customField: record.customField ?? null
  });
}

await browser.close();

Replace the URL and predicate with the route you observed. If several requests match, also check the method, query parameters, response status and a distinctive response header or URL segment. Validate the payload shape before processing it; a login page or an error object can otherwise be mistaken for an empty result.

Capture a request caused by a control

For a tab or “load more” button, create the promise first and await it alongside the action:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const responsePromise = page.waitForResponse(r =>
  r.url().includes('/api/records') && r.request().method() === 'GET'
);
await page.getByRole('button', { name: 'Load more' }).click();
const response = await responsePromise;
const nextPage = await response.json();

This ordering prevents a fast response from being missed. For APIs that return cursors or next links, persist the cursor from each payload and stop only when the server indicates there is no next page.

When the value exists only in the rendered DOM

Use semantic roles, labels and stable data-* attributes rather than generated CSS classes. Scope every selector to one record so a duplicate label in a sidebar or another card cannot contaminate the result.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();

await page.goto('https://example.com/profile/123', {
  waitUntil: 'domcontentloaded'
});

const card = page.locator('[data-record-id="123"]');
await card.getByRole('button', { name: 'Details' }).click();
const field = card.locator('[data-field="customer-tier"]');
await field.waitFor({ state: 'visible' });

const value = (await field.textContent())?.trim() ?? null;
console.log({ recordId: '123', customField: value });

await browser.close();

If the field is represented by an attribute, read it explicitly with getAttribute(). If a component virtualizes rows, scroll the record into view before waiting. For values formatted by JavaScript, the DOM is the authoritative presentation, while the network payload may contain a different raw representation.

Waiting correctly: state, not sleep

Fixed sleeps are slow when the page is fast and flaky when it is slow. Prefer one of these completion conditions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • A response with the expected URL, method and status.
  • A locator becoming visible or containing non-empty text.
  • A record count reaching the expected value after pagination.
  • A known loading indicator disappearing, combined with a field-specific check.

Set explicit navigation and response timeouts, and log the URL, record ID, status and elapsed time. A timeout should produce a replayable failed-record entry, not silently discard data.

Pagination, scrolling and retries

Cursor or next-link pagination

Read the cursor or next URL from each JSON response, request the next page through the application’s normal interaction when possible, and persist progress after every successful page. Deduplicate by a stable record ID because a changing dataset can overlap pages.

Infinite scroll

Scroll the list container, then wait for either the specific response or an increase in rendered records. Stop when the API returns an empty page, a terminal cursor, or the UI exposes no further loading control. Do not assume a fixed number of scrolls.

Retry policy

Retry transient network failures and 5xx responses with a small capped exponential backoff. Do not blindly retry authentication failures, validation errors or a stable 4xx response. Store the original URL, request parameters and error so the record can be replayed without rerunning the entire crawl.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Authentication, context and browser behavior

Create the browser context with the cookies, storage state and headers required by the site. A context keeps those credentials isolated from other jobs. If a flow requires a login, automate it once, save the authenticated storage state according to the site’s rules, and load it for subsequent runs rather than logging in for every record.

Service workers can change where requests originate. Playwright notes that page-level routing does not intercept service-worker requests. If an expected call is missing, inspect service-worker activity and consider a context-level route or a context configured to block service workers when that is appropriate for the target.

Selenium as an alternative

Selenium’s JavaScript API installs with npm install selenium-webdriver. Selenium Manager can handle browser-driver installation, and Selenium supports simulated user actions and arbitrary JavaScript execution. A minimal DOM extraction looks like this:

const { Builder, By } = require('selenium-webdriver');

(async function scrape() {
  const driver = await new Builder().forBrowser('chrome').build();
  try {
    await driver.get('https://example.com/profile/123');
    const details = await driver.findElement(
      By.css('[data-field="customer-tier"]')
    );
    console.log(await details.getText());
  } finally {
    await driver.quit();
  }
})();

Choose between Playwright and Selenium based on browser coverage, network-interception ergonomics, locator quality, team language, hosting cost and observability. Whichever tool you choose, retain the same semantic waits, record scoping and retry rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hosted rendering when you do not want to run a browser

Cloudflare’s Browser Run /content endpoint navigates to a URL and returns fully rendered HTML, including the head section, after JavaScript execution. It can simplify deployment for JavaScript-heavy pages, but verify authentication handling, quotas, cost and the target’s terms for your use case. You still need a parser, pagination logic and field-specific validation after receiving the HTML.

Normalize and audit the extracted values

  • Keep null distinct from a missing property; “not present” and “present but empty” can have different meanings.
  • Flatten nested objects deliberately, documenting how arrays and localized values are represented.
  • Store the source URL, record ID, extraction timestamp, response status and scraper version with each result.
  • Preserve raw payloads or hashes when policy permits so a changed value can be audited.
  • Validate types and expected ranges, and flag an unexpected schema instead of coercing it silently.

Common failures and precise fixes

HTML is empty or contains only a shell

The request did not execute the SPA. Use Playwright or Selenium, confirm the correct route, and wait for the field-specific locator or API response.

The expected response was never captured

The listener was registered too late, the URL predicate is too broad or the request came from a service worker. Create the promise before the action, tighten the predicate, and inspect context-level network events.

The field appears only after scrolling

Scroll the correct container, then await the response or locator caused by that scroll. Virtualized lists may remove off-screen nodes, so extract each record while it is mounted or use the underlying API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A selector broke after a redesign

Replace generated classes with a role, accessible label, stable attribute or a documented test identifier. Keep the selector scoped to the record container.

Values are duplicated or stale

Verify the record ID in both the payload and the DOM container. Wait for the loading state to finish and avoid selecting the first matching label on the page.

Pagination has gaps

Persist every cursor or next link, record response statuses and deduplicate by ID. Re-run only failed pages from the saved checkpoint.

A bot check or login page is returned

Stop and handle the site’s permitted authentication or access process; do not attempt to bypass a challenge. Treat the response as a failed extraction and log it for review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an extraction layer with this decision table

Need Best first choice Reason
Structured custom field in JSON Network/API capture Less coupled to presentation markup and easier to validate.
Computed or interaction-only value Rendered DOM Reads what the user actually sees after scripts and formatting run.
Complex browser interaction Playwright or Selenium Provides navigation, actions, waits and authenticated context.
Managed rendered HTML Cloudflare Browser Run Removes local browser hosting, subject to deployment-specific limits.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your immediate need is a clean visual capture of the rendered page for QA, archives or an AI workflow, ScreenshotNeo provides a website screenshot API and MCP server. It is not a replacement for parsing JSON custom fields, but it can capture the final page without you maintaining a browser. Before capture it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

One GET request returns PNG, JPEG, WebP or PDF. See the ScreenshotNeo documentation for all options, including full-page lazy-image loading, CSS-selector element capture, device and retina settings, custom CSS or JavaScript, click and wait conditions, request blocking, cookies and headers, timezone and geolocation, caching, signed links, asynchronous webhooks, bulk capture and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://example.com/profile/123 
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/profile/123"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/profile/123'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try a capture.

Operational cost and reliability notes

Browser automation consumes CPU and memory and is slower than replaying a documented JSON endpoint. Reuse a browser process where safe, limit concurrency to what the target permits, block unnecessary resources only when it does not alter the field, and cache immutable pages. Keep timeouts finite and expose metrics for success, timeout, HTTP status, records per page and retry count. A scraper that returns fewer records without an error is more dangerous than one that fails loudly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I scrape a JavaScript SPA with only curl?

Only if you reproduce the SPA’s underlying authenticated API calls yourself. Curl does not execute the page’s JavaScript, so it will normally receive the application shell.

How do I know whether a custom field is in the API response?

Inspect Fetch/XHR traffic while revealing the field, then search the response payload for its label, record ID or value. If it is absent and appears only after client-side computation, use a scoped DOM locator.

What should I store to make a failed scrape reproducible?

Store the source URL, record ID, request parameters or cursor, response status, extraction timestamp, scraper version and the failure message; retain raw data when your policy allows.

Is a rendered screenshot evidence that the field was extracted correctly?

No. A screenshot verifies visual output only. Validate the structured value from the API or DOM and keep record IDs and response metadata for an auditable extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.