Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk7 min

How to Scrape Dynamic Websites with JavaScript

A practical JavaScript guide to inspecting dynamic pages, choosing request-based extraction or browser automation, and validating what you collect.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape a dynamic website, first check whether the data is available from a repeatable network request; if it is, reproduce that request instead of rendering the whole page. Use JavaScript browser automation when the page’s JavaScript, interaction, or displayed browser state is essential. This guide shows both approaches, with Playwright examples for browser-driven extraction.

Choose the extraction method before writing a scraper

A page can look empty in its initial HTML and fill in after JavaScript runs. That does not automatically mean you need a browser. Inspect the page’s network activity and find the request that supplies the data. Scrapy’s guide to dynamic content says reproducing the request containing the desired data is preferred when practical: it can return structured data with less parsing and network transfer than rendering a page. Scrapy: Selecting dynamically-loaded content.

  • Use a direct request when a repeatable request provides the fields you need and you can responsibly reproduce it.
  • Use browser automation when the relevant request is difficult to reproduce, results depend on page state or interaction, or you need what the browser actually displays.
  • Consider a managed browser when operating browser instances or coordinating a site-wide crawl is an infrastructure requirement, rather than a prerequisite for a small scrape.

These are different methods, not a ranking of universal speed or reliability. The official documentation reviewed here establishes capabilities, not comparative benchmarks.

Inspect the page and identify when its data arrives

  1. Open the page in a browser. Note which content is missing initially and what action or delay makes it appear.
  2. Inspect network requests. Look for a request whose response contains the desired fields. Check whether it is repeatable and whether its use is appropriate for your purpose.
  3. Choose the smallest method that meets the need. A structured response may be easier to validate than text parsed from rendered markup. If you need the rendered result or interaction, continue with a browser.
  4. Test one page and a small sample first. Compare extracted values with the page, account for absent or changed fields, and record the source URL and retrieval time with your data.

Those validation practices are practical safeguards; the cited documentation does not define a universal schema or validation protocol.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Playwright when the browser must run the page

Playwright’s Page API supports observing and routing requests and waiting for page events or selector conditions. Prefer waits tied to evidence that the page is ready over a fixed pause. Playwright Page API.

Install and run a minimal JavaScript scraper

The following example uses Node.js with Playwright’s library. Replace the sample URL and selector with the page and element you are authorized to access. The selector is deliberately site-specific: identify it with the browser’s developer tools before running the script.

  1. In a new project directory, run npm init -y.
  2. Install Playwright with npm install playwright, then install its browser with npx playwright install chromium.
  3. Save this as scrape.js and replace https://example.com/catalog and .product-card with the target page and its repeating item selector.
  4. Run node scrape.js. The script prints JSON to the terminal.
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage();

  try {
    await page.goto('https://example.com/catalog', {
      waitUntil: 'domcontentloaded',
      timeout: 30000,
    });

    // Wait for actual page content, not an arbitrary sleep.
    await page.locator('.product-card').first().waitFor({
      state: 'visible',
      timeout: 15000,
    });

    const products = await page.locator('.product-card').evaluateAll(cards =>
      cards.map(card => ({
        name: card.querySelector('.product-name')?.textContent?.trim() ?? null,
        href: card.querySelector('a')?.href ?? null,
      }))
    );

    console.log(JSON.stringify(products, null, 2));
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

The locator wait fails with a timeout if no matching visible card appears in the allotted time; it does not prove that every item has loaded. For pages that append results as you scroll, implement and verify the site’s actual pagination or scrolling behavior rather than assuming one DOM snapshot is complete. For interaction-dependent content, use locators to perform the required action and then wait for a resulting locator, URL, or response.

Use a response instead of DOM parsing when it is the better source

If inspection reveals a request containing the needed data, you can observe matching responses in Playwright and inspect their bodies. Match the request narrowly; a page may make many unrelated requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const responsePromise = page.waitForResponse(response =>
  response.url().includes('/api/catalog') && response.status() === 200
);

await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
const response = await responsePromise;
const data = await response.json();
console.log(JSON.stringify(data, null, 2));

This snippet assumes the page triggers a JSON response at a URL containing /api/catalog. Replace that match with the observed request pattern and adjust parsing if the response is not JSON. Playwright also supports request observation and routing through its Page API.

Puppeteer is another browser-automation option

Puppeteer is suitable for browser-driven interaction in JavaScript. Its documentation recommends locator-based interaction; locators wait for an element to exist and be ready for the action. Use the official Puppeteer page interactions guide for current API details. This guide uses Playwright for runnable examples rather than suggesting both libraries are required.

Scale only after a single-page extraction is dependable

For multiple pages, first establish that the extraction works for representative pages and that your waits, parsing, and failure handling match the site. Keep browser control and the expected request volume proportionate to the task. Cloudflare Browser Run documents several distinct hosted options: Quick Actions for simple scrape tasks, browser sessions controlled through Playwright, Puppeteer, CDP, or Stagehand, and a crawl endpoint for site-wide extraction. Its documentation says the crawl endpoint returns asynchronous results and describes availability on Free and Paid plans; check the current documentation for terms and availability before relying on them. Cloudflare Browser Run (page last updated August 11, 2026, according to Cloudflare).

A managed service is an infrastructure choice, not a requirement to use JavaScript or to scrape a small number of pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrape responsibly

Google says its automated crawlers use the Robots Exclusion Protocol and download and parse robots.txt before crawling. Google also explains that a file’s rules apply to the host, protocol, and port where that file is served. These statements describe Google’s crawler guidance, not a complete rulebook for every scraper. Google: robots.txt specifications.

Before scraping a particular site, separately review its terms, access controls, privacy implications, applicable law, and your intended use of the data. A robots.txt rule alone does not determine whether a particular scrape is permitted.

Troubleshoot common failures

  • The selector times out: confirm the selector against the live DOM and check whether the content appears only after an interaction, a response, or scrolling. Wait for the meaningful condition rather than extending a fixed sleep without evidence.
  • The page loads but extracted fields are null: inspect the item markup; the site may have changed its structure, or the field may be absent for some items. Make parsing tolerant of missing values and validate a sample.
  • The browser shows fewer results than expected: determine whether the page paginates or loads more items on scroll. A visible first batch is not proof that the whole collection is present.
  • The response wait never resolves: verify the page actually triggers the request, and refine the URL and status match using observed network activity. The example’s /api/catalog pattern is illustrative, not a universal endpoint.
  • Navigation exceeds the timeout: distinguish a navigation failure from a page that loaded its initial document but is still fetching content. Use a navigation condition and a separate wait for the specific content needed.
  • A managed crawl behaves differently from a direct session: check which Cloudflare Browser Run option you selected; Quick Actions, browser sessions, and the crawl endpoint serve distinct use cases.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a screenshot or PDF of a dynamic page rather than structured field extraction, ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. Its clean-shot steps accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients.

For example, install Python’s requests package and run:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

See the ScreenshotNeo API documentation for authentication and request options. This captures a page image; it is not a substitute for extracting structured records from a site’s data request or DOM.

ScreenshotNeo’s Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

Frequently Asked Questions

Can a scraper retrieve data from a page that requires JavaScript?

Yes. Reproduce the request that supplies the data when practical, or use browser automation to run the page when its JavaScript or interaction is necessary.

Does robots.txt grant permission to scrape a website?

No. It communicates crawler rules within its scope; review the site’s terms, access controls, privacy concerns, applicable law, and intended data use separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.