Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A browser can show data that a basic scraper never receives because the two are looking at different things. A scraper built with an HTTP client such as requests or Scrapy typically downloads the server’s initial HTML response. A modern page may send only an app shell in that response, then use JavaScript to fetch data and add it to the live page. The browser’s rendered DOM can contain that data even though the original response does not.

That distinction also explains why a selector can work in Inspect Element but return nothing from your scraper: the selector may be fine; the data may simply not be in the HTML your code downloaded.

View Source and Inspect Element show different versions of a page

View Source shows the HTML response the server sent. Inspect Element shows the browser’s current DOM, which scripts may have changed after receiving that response. The browser can execute JavaScript, make additional requests, and insert returned values into the page. The live DOM is therefore not proof that those values were present in the original HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google Search Central describes this app-shell pattern: some JavaScript sites send initial HTML without the actual content and require JavaScript execution to generate it. The practical test is to save the exact response your scraper received and search it for the visible text or data. If it is absent there, changing CSS selectors alone will not make it appear.

How to compare the response with the rendered page

  1. Request the page with the same URL and relevant headers your scraper uses; save or print the response body.
  2. Compare that response with the browser’s View Source, not only with the Elements panel.
  3. Search the response for a distinctive value visible on screen, such as a table cell or result title.
  4. If the value is missing, investigate a follow-up request, JavaScript execution, or browser state before adjusting the extraction selector.

A concise first check with Python is:

import requests

url = "https://example.com/"
response = requests.get(url, timeout=30)
print("HTTP status:", response.status_code)
print(response.text[:5000])

Replace the example URL with a page you are authorized to access. This prints the response returned to the HTTP client; it does not execute the page’s scripts.

Find the request that supplies the missing data

When content is absent from the initial response, the next question is often not “How do I render this page?” but “Which request brought this data to the browser?” Many pages retrieve results with XHR or fetch, often as JSON, and then build the visible table or cards from that response.

  1. Open the page in a browser and open Developer Tools, then select the Network panel.
  2. Filter for XHR/fetch or look for JSON, GraphQL, and document requests.
  3. Clear the request list, then reload the page or repeat the action that reveals the data. For pagination, search, or “load more,” trigger that action too.
  4. Select likely requests and inspect the URL, method, query parameters, request body, headers, cookies, and response. Look for the actual values you need rather than relying on request names.
  5. Try reproducing the request directly using an HTTP client. Preserve the relevant method, parameters, body, and permitted authentication state; compare its response with the browser’s response.

Scrapy’s guidance is to first download the page with an HTTP client such as curl or wget and check whether the information is in that response. If it is not, use browser network tools to locate the request that carries it. A direct data request is usually the leaner approach when it works and is permitted: it avoids rendering a whole page and returns structured data that is easier to process.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why copying a request may stop working

A request observed once is not necessarily a stable public interface. It may depend on a short-lived token, a session cookie, a changing query value, or a sequence of earlier requests. Reproduce the necessary flow rather than blindly copying a stale request. If the site requires authentication, use valid credentials and permission; do not try to defeat access controls.

When to use a browser instead of a direct request

Use browser automation when the site genuinely requires JavaScript execution or an interaction to make the data available—for example, a click that triggers the request, scrolling that loads more results, browser storage, an iframe, or a complex permitted login flow. Playwright browser contexts support JavaScript and authentication settings, and its network APIs can observe requests and wait for a response caused by an action.

Here is a Node.js example that opens a page, clicks a button, waits for a response from that interaction, and prints the response body. Change the URL, button selector, and response URL check to match the page you are permitted to automate. Install Playwright with npm install playwright and install its browser with npx playwright install chromium.

const { chromium } = require("playwright");

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage();

  try {
    await page.goto("https://example.com/", { waitUntil: "domcontentloaded" });
    const responsePromise = page.waitForResponse(response =>
      response.url().includes("/api/") && response.ok()
    );
    await page.locator("button").first().click();
    const response = await responsePromise;
    console.log("Data response URL:", response.url());
    console.log(await response.text());
  } finally {
    await browser.close();
  }
})();

The response URL condition is deliberately generic; make it specific enough to match the request that actually contains the data. If the page makes several API calls, a broad condition may wait for the wrong one. If there is no click to perform, remove the click and wait for the relevant request during navigation instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract from the response or from the DOM?

  • Prefer the response payload when it contains the needed records. It is typically more structured and avoids fragile selectors tied to page layout.
  • Read the rendered DOM when the data is only exposed after browser-side transformations or when the page itself is the required output.
  • Use a browser interaction when the response is triggered by a click, scroll, or other user action and cannot be reproduced as a stable, permitted direct request.

Wait for the data, not just for the page

A navigation event does not guarantee that asynchronous results have arrived. The document may be loaded while a client-side request is still running, so extracting immediately can produce an empty table or incomplete results.

Wait for a condition tied to the content you need: a specific successful response, a selector that contains real data, or an appropriate network-idle state. Playwright supports waiting for a specific response, and Cloudflare’s browser-rendering guidance describes network-idle waits for rendered extraction. Avoid relying on a short fixed sleep as the main readiness check: network speed and page behavior vary, and a delay can be both too short on a slow run and wastefully long on a fast one.

Check browser state, service workers, and cross-origin access

Cookies, credentials, and headers

The browser may have cookies, HTTP credentials, or headers that the scraper does not. A page can therefore return different content to the two clients. When reproducing a request, identify which state is legitimately required and configure it deliberately. Playwright contexts support credentials and authentication settings; do not expose secrets in logs or commit them to source control.

Service workers

A service worker can intercept requests. Playwright documents that service workers may make some network events unavailable to its routing APIs; if expected requests are missing from your instrumentation, check whether a service worker is involved and whether blocking service workers for the context is appropriate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CORS is not a general scraper block

CORS governs whether browser JavaScript is allowed to read certain cross-origin responses. MDN notes that a no-cors response is opaque to JavaScript. This is different from an HTTP client downloading a response: CORS is a browser security mechanism, not a universal rule preventing a server-side client from making a request. Do not treat disabling browser protections as a valid substitute for authorization or permission.

Choose the lightest method that meets the need

Approach Best fit Trade-off
Direct HTTP or API request A permitted endpoint reliably returns the needed data and can be called with the required request details. Fast and light, but can depend on session state, tokens, or an endpoint that changes.
Browser automation, such as Playwright Data requires JavaScript, interaction, browser storage, or rendered-page behavior. Closer to the user-visible flow, but uses more resources and requires maintaining browser steps and selectors.
Managed browser rendering You need hosted browser execution or rendered HTML or element extraction rather than running browsers yourself. Can reduce local browser operations, but introduces a hosted service into the workflow.

Cloudflare documents browser rendering options for element scraping and a fully rendered HTML content endpoint. Choose a managed option only if its capabilities and terms fit the job; it does not remove the need to identify the right data or respect the site’s access rules.

Or skip the browser setup

If the job is to capture what a page looks like—not to retrieve structured records for a data pipeline—ScreenshotNeo can return a screenshot or PDF through one GET request. It is a screenshot API and MCP server, not a replacement for extracting JSON from an endpoint. For example, the following saves a WebP capture of Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo’s free plan to try it without a card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot an empty or incomplete scrape

  • The response contains no target text: inspect Network for the request that supplies it, then test that request directly.
  • The API response is empty or different from the browser: compare method, URL, query/body, headers, cookies, and authentication state. Check whether a required token or session expires.
  • The request succeeds but the extracted results are empty: verify that you are waiting for the correct response or for a selector that contains data, not merely for document navigation.
  • Only results after a click or scroll are missing: reproduce that action in the browser, or identify and call the request it triggers if it is stable and permitted.
  • Network instrumentation shows no expected request: check filters and request timing, then investigate whether a service worker intercepts it.
  • Browser JavaScript cannot read a cross-origin response: inspect the browser’s CORS error and the response’s access-control behavior; do not assume that a no-cors response exposes readable data.
  • The page works only when signed in: use authorized credentials or stored authentication state and confirm that your access permits automation.
  • The page shows a challenge or blocks automation: stop and check the site’s rules and contact the site owner if needed. Do not bypass a CAPTCHA or other access control.

Reliability, performance, and responsible collection

Direct requests generally avoid the CPU and startup cost of launching a browser, but they are only reliable while the endpoint and required state remain usable. Browser automation follows more of the visible experience, at the cost of additional runtime and maintenance when page behavior changes. For either approach, wait on meaningful conditions, handle request failures explicitly, and avoid collecting more often than the site permits.

Before automating, check the site’s authorization requirements, robots directives, terms, rate limits, and privacy obligations. A value visible in a browser is not, by itself, permission to collect it. Use valid credentials for protected content and do not defeat access controls.

Frequently Asked Questions

Can I scrape data from a page that requires JavaScript?

Yes, when you are authorized to collect it. First look for a permitted request that provides the data; use a browser when JavaScript execution or interaction is genuinely required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does the same CSS selector work in DevTools but not in my scraper?

The browser’s live DOM may include elements inserted after the initial HTML response. Check the response your scraper received before treating the selector as the cause.

Does CORS prevent Python requests or Scrapy from downloading a page?

CORS is enforced by browsers for cross-origin access by scripts; it is not a universal restriction on server-side HTTP clients. Authentication, server rules, and permission still apply.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.