The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use a real browser to run the SPA, wait for the field’s populated state, and then extract either the API response that supplied it or the rendered DOM. A normal HTTP request often returns only an application shell. The most reliable workflow is to discover the request carrying the custom field, capture it before the interaction that triggers it, and parse its JSON. If the value exists only after client-side formatting or an interaction, scope a locator to the correct record and read the rendered element.
Why a normal HTTP request misses custom fields
React, Vue and Angular applications commonly deliver a small HTML shell plus JavaScript bundles. The browser then requests records, applies authentication and state, renders components, and reveals fields after a tab click, search, scroll or “load more” action. An HTTP client such as requests or curl sees the shell unless you separately reproduce those data calls.
There are two useful extraction layers:
- Network/API layer: capture the JSON response containing the record and custom field. This is usually less brittle than selectors tied to presentation markup.
- Rendered-DOM layer: let the browser finish rendering, then read text, attributes or links from a stable element. Use this when the value is computed in the browser or is not present in a usable response.
Check the site’s robots directives, terms, authentication requirements, privacy and copyright obligations, rate limits and applicable law before collecting data.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesMap the page state before writing a scraper
- Identify the route and record key. Record the URL pattern, record ID and the container that represents one record.
- Find the field’s label and state. Note whether the custom field is visible immediately or appears after a Details tab, search, pagination, “load more” control or scrolling.
- Observe the browser’s requests. In developer tools, filter Fetch/XHR calls while performing the interaction. Look for JSON containing the record ID or field name.
- Record the trigger and completion signal. The trigger may be navigation or a click; completion should be a known response or a locator becoming visible, not an arbitrary delay.
Keep authentication cookies, local storage and headers in the same browser context that performs navigation. Otherwise the page can look logged out even when your interactive browser is authenticated.
#1 Best Overall
Preferred method: capture the SPA’s JSON response
Register the response listener before navigation or the click that causes the request. The following Playwright script waits for a GET to /api/records, parses the payload, and preserves a null value when a field is absent.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/records') &&
response.request().method() === 'GET' &&
response.status() === 200
);
await page.goto('https://example.com/records', { waitUntil: 'domcontentloaded' });
const response = await responsePromise;
const payload = await response.json();
for (const record of payload.records ?? []) {
console.log({
id: record.id,
customField: record.customField ?? null
});
}
await browser.close();
Replace the URL and predicate with the route you observed. If several requests match, also check the method, query parameters, response status and a distinctive response header or URL segment. Validate the payload shape before processing it; a login page or an error object can otherwise be mistaken for an empty result.
Capture a request caused by a control
For a tab or “load more” button, create the promise first and await it alongside the action:
const responsePromise = page.waitForResponse(r =>
r.url().includes('/api/records') && r.request().method() === 'GET'
);
await page.getByRole('button', { name: 'Load more' }).click();
const response = await responsePromise;
const nextPage = await response.json();
This ordering prevents a fast response from being missed. For APIs that return cursors or next links, persist the cursor from each payload and stop only when the server indicates there is no next page.
When the value exists only in the rendered DOM
Use semantic roles, labels and stable data-* attributes rather than generated CSS classes. Scope every selector to one record so a duplicate label in a sidebar or another card cannot contaminate the result.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto('https://example.com/profile/123', {
waitUntil: 'domcontentloaded'
});
const card = page.locator('[data-record-id="123"]');
await card.getByRole('button', { name: 'Details' }).click();
const field = card.locator('[data-field="customer-tier"]');
await field.waitFor({ state: 'visible' });
const value = (await field.textContent())?.trim() ?? null;
console.log({ recordId: '123', customField: value });
await browser.close();
If the field is represented by an attribute, read it explicitly with getAttribute(). If a component virtualizes rows, scroll the record into view before waiting. For values formatted by JavaScript, the DOM is the authoritative presentation, while the network payload may contain a different raw representation.
Waiting correctly: state, not sleep
Fixed sleeps are slow when the page is fast and flaky when it is slow. Prefer one of these completion conditions:
- A response with the expected URL, method and status.
- A locator becoming visible or containing non-empty text.
- A record count reaching the expected value after pagination.
- A known loading indicator disappearing, combined with a field-specific check.
Set explicit navigation and response timeouts, and log the URL, record ID, status and elapsed time. A timeout should produce a replayable failed-record entry, not silently discard data.
Pagination, scrolling and retries
Cursor or next-link pagination
Read the cursor or next URL from each JSON response, request the next page through the application’s normal interaction when possible, and persist progress after every successful page. Deduplicate by a stable record ID because a changing dataset can overlap pages.
Infinite scroll
Scroll the list container, then wait for either the specific response or an increase in rendered records. Stop when the API returns an empty page, a terminal cursor, or the UI exposes no further loading control. Do not assume a fixed number of scrolls.
Retry policy
Retry transient network failures and 5xx responses with a small capped exponential backoff. Do not blindly retry authentication failures, validation errors or a stable 4xx response. Store the original URL, request parameters and error so the record can be replayed without rerunning the entire crawl.
Free tools Windows power users keep installed
One-click scans. No signup required.
Authentication, context and browser behavior
Create the browser context with the cookies, storage state and headers required by the site. A context keeps those credentials isolated from other jobs. If a flow requires a login, automate it once, save the authenticated storage state according to the site’s rules, and load it for subsequent runs rather than logging in for every record.
Rank #3
Service workers can change where requests originate. Playwright notes that page-level routing does not intercept service-worker requests. If an expected call is missing, inspect service-worker activity and consider a context-level route or a context configured to block service workers when that is appropriate for the target.
Selenium as an alternative
Selenium’s JavaScript API installs with npm install selenium-webdriver. Selenium Manager can handle browser-driver installation, and Selenium supports simulated user actions and arbitrary JavaScript execution. A minimal DOM extraction looks like this:
const { Builder, By } = require('selenium-webdriver');
(async function scrape() {
const driver = await new Builder().forBrowser('chrome').build();
try {
await driver.get('https://example.com/profile/123');
const details = await driver.findElement(
By.css('[data-field="customer-tier"]')
);
console.log(await details.getText());
} finally {
await driver.quit();
}
})();
Choose between Playwright and Selenium based on browser coverage, network-interception ergonomics, locator quality, team language, hosting cost and observability. Whichever tool you choose, retain the same semantic waits, record scoping and retry rules.
Hosted rendering when you do not want to run a browser
Cloudflare’s Browser Run /content endpoint navigates to a URL and returns fully rendered HTML, including the head section, after JavaScript execution. It can simplify deployment for JavaScript-heavy pages, but verify authentication handling, quotas, cost and the target’s terms for your use case. You still need a parser, pagination logic and field-specific validation after receiving the HTML.
Normalize and audit the extracted values
- Keep
nulldistinct from a missing property; “not present” and “present but empty” can have different meanings. - Flatten nested objects deliberately, documenting how arrays and localized values are represented.
- Store the source URL, record ID, extraction timestamp, response status and scraper version with each result.
- Preserve raw payloads or hashes when policy permits so a changed value can be audited.
- Validate types and expected ranges, and flag an unexpected schema instead of coercing it silently.
Common failures and precise fixes
HTML is empty or contains only a shell
The request did not execute the SPA. Use Playwright or Selenium, confirm the correct route, and wait for the field-specific locator or API response.
The expected response was never captured
The listener was registered too late, the URL predicate is too broad or the request came from a service worker. Create the promise before the action, tighten the predicate, and inspect context-level network events.
Rank #4
The field appears only after scrolling
Scroll the correct container, then await the response or locator caused by that scroll. Virtualized lists may remove off-screen nodes, so extract each record while it is mounted or use the underlying API.
A selector broke after a redesign
Replace generated classes with a role, accessible label, stable attribute or a documented test identifier. Keep the selector scoped to the record container.
Values are duplicated or stale
Verify the record ID in both the payload and the DOM container. Wait for the loading state to finish and avoid selecting the first matching label on the page.
Pagination has gaps
Persist every cursor or next link, record response statuses and deduplicate by ID. Re-run only failed pages from the saved checkpoint.
A bot check or login page is returned
Stop and handle the site’s permitted authentication or access process; do not attempt to bypass a challenge. Treat the response as a failed extraction and log it for review.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Choose an extraction layer with this decision table
| Need | Best first choice | Reason |
|---|---|---|
| Structured custom field in JSON | Network/API capture | Less coupled to presentation markup and easier to validate. |
| Computed or interaction-only value | Rendered DOM | Reads what the user actually sees after scripts and formatting run. |
| Complex browser interaction | Playwright or Selenium | Provides navigation, actions, waits and authenticated context. |
| Managed rendered HTML | Cloudflare Browser Run | Removes local browser hosting, subject to deployment-specific limits. |
Or skip the browser setup
If your immediate need is a clean visual capture of the rendered page for QA, archives or an AI workflow, ScreenshotNeo provides a website screenshot API and MCP server. It is not a replacement for parsing JSON custom fields, but it can capture the final page without you maintaining a browser. Before capture it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
One GET request returns PNG, JPEG, WebP or PDF. See the ScreenshotNeo documentation for all options, including full-page lazy-image loading, CSS-selector element capture, device and retina settings, custom CSS or JavaScript, click and wait conditions, request blocking, cookies and headers, timezone and geolocation, caching, signed links, asynchronous webhooks, bulk capture and usage reporting.
Best Value
curl -G "https://api.screenshotneo.com/v1/shot"
-d access_key=YOUR_API_KEY
--data-urlencode url=https://example.com/profile/123
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/profile/123"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/profile/123'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try a capture.
Operational cost and reliability notes
Browser automation consumes CPU and memory and is slower than replaying a documented JSON endpoint. Reuse a browser process where safe, limit concurrency to what the target permits, block unnecessary resources only when it does not alter the field, and cache immutable pages. Keep timeouts finite and expose metrics for success, timeout, HTTP status, records per page and retry count. A scraper that returns fewer records without an error is more dangerous than one that fails loudly.
Recommended Free Tools
Frequently Asked Questions
Can I scrape a JavaScript SPA with only curl?
Only if you reproduce the SPA’s underlying authenticated API calls yourself. Curl does not execute the page’s JavaScript, so it will normally receive the application shell.
How do I know whether a custom field is in the API response?
Inspect Fetch/XHR traffic while revealing the field, then search the response payload for its label, record ID or value. If it is absent and appears only after client-side computation, use a scoped DOM locator.
What should I store to make a failed scrape reproducible?
Store the source URL, record ID, request parameters or cursor, response status, extraction timestamp, scraper version and the failure message; retain raw data when your policy allows.
Is a rendered screenshot evidence that the field was extracted correctly?
No. A screenshot verifies visual output only. Validate the structured value from the API or DOM and keep record IDs and response metadata for an auditable extraction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

