To scrape a dynamic website, first check whether the data is available from a repeatable network request; if it is, reproduce that request instead of rendering the whole page. Use JavaScript browser automation when the page’s JavaScript, interaction, or displayed browser state is essential. This guide shows both approaches, with Playwright examples for browser-driven extraction.
Choose the extraction method before writing a scraper
A page can look empty in its initial HTML and fill in after JavaScript runs. That does not automatically mean you need a browser. Inspect the page’s network activity and find the request that supplies the data. Scrapy’s guide to dynamic content says reproducing the request containing the desired data is preferred when practical: it can return structured data with less parsing and network transfer than rendering a page. Scrapy: Selecting dynamically-loaded content.
- Use a direct request when a repeatable request provides the fields you need and you can responsibly reproduce it.
- Use browser automation when the relevant request is difficult to reproduce, results depend on page state or interaction, or you need what the browser actually displays.
- Consider a managed browser when operating browser instances or coordinating a site-wide crawl is an infrastructure requirement, rather than a prerequisite for a small scrape.
These are different methods, not a ranking of universal speed or reliability. The official documentation reviewed here establishes capabilities, not comparative benchmarks.
Inspect the page and identify when its data arrives
- Open the page in a browser. Note which content is missing initially and what action or delay makes it appear.
- Inspect network requests. Look for a request whose response contains the desired fields. Check whether it is repeatable and whether its use is appropriate for your purpose.
- Choose the smallest method that meets the need. A structured response may be easier to validate than text parsed from rendered markup. If you need the rendered result or interaction, continue with a browser.
- Test one page and a small sample first. Compare extracted values with the page, account for absent or changed fields, and record the source URL and retrieval time with your data.
Those validation practices are practical safeguards; the cited documentation does not define a universal schema or validation protocol.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Use Playwright when the browser must run the page
Playwright’s Page API supports observing and routing requests and waiting for page events or selector conditions. Prefer waits tied to evidence that the page is ready over a fixed pause. Playwright Page API.
Install and run a minimal JavaScript scraper
The following example uses Node.js with Playwright’s library. Replace the sample URL and selector with the page and element you are authorized to access. The selector is deliberately site-specific: identify it with the browser’s developer tools before running the script.
- In a new project directory, run
npm init -y. - Install Playwright with
npm install playwright, then install its browser withnpx playwright install chromium. - Save this as
scrape.jsand replacehttps://example.com/catalogand.product-cardwith the target page and its repeating item selector. - Run
node scrape.js. The script prints JSON to the terminal.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
try {
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 30000,
});
// Wait for actual page content, not an arbitrary sleep.
await page.locator('.product-card').first().waitFor({
state: 'visible',
timeout: 15000,
});
const products = await page.locator('.product-card').evaluateAll(cards =>
cards.map(card => ({
name: card.querySelector('.product-name')?.textContent?.trim() ?? null,
href: card.querySelector('a')?.href ?? null,
}))
);
console.log(JSON.stringify(products, null, 2));
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
The locator wait fails with a timeout if no matching visible card appears in the allotted time; it does not prove that every item has loaded. For pages that append results as you scroll, implement and verify the site’s actual pagination or scrolling behavior rather than assuming one DOM snapshot is complete. For interaction-dependent content, use locators to perform the required action and then wait for a resulting locator, URL, or response.
Rank #2
Use a response instead of DOM parsing when it is the better source
If inspection reveals a request containing the needed data, you can observe matching responses in Playwright and inspect their bodies. Match the request narrowly; a page may make many unrelated requests.
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/catalog') && response.status() === 200
);
await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
const response = await responsePromise;
const data = await response.json();
console.log(JSON.stringify(data, null, 2));
This snippet assumes the page triggers a JSON response at a URL containing /api/catalog. Replace that match with the observed request pattern and adjust parsing if the response is not JSON. Playwright also supports request observation and routing through its Page API.
Puppeteer is another browser-automation option
Puppeteer is suitable for browser-driven interaction in JavaScript. Its documentation recommends locator-based interaction; locators wait for an element to exist and be ready for the action. Use the official Puppeteer page interactions guide for current API details. This guide uses Playwright for runnable examples rather than suggesting both libraries are required.
Scale only after a single-page extraction is dependable
For multiple pages, first establish that the extraction works for representative pages and that your waits, parsing, and failure handling match the site. Keep browser control and the expected request volume proportionate to the task. Cloudflare Browser Run documents several distinct hosted options: Quick Actions for simple scrape tasks, browser sessions controlled through Playwright, Puppeteer, CDP, or Stagehand, and a crawl endpoint for site-wide extraction. Its documentation says the crawl endpoint returns asynchronous results and describes availability on Free and Paid plans; check the current documentation for terms and availability before relying on them. Cloudflare Browser Run (page last updated August 11, 2026, according to Cloudflare).
A managed service is an infrastructure choice, not a requirement to use JavaScript or to scrape a small number of pages.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsScrape responsibly
Google says its automated crawlers use the Robots Exclusion Protocol and download and parse robots.txt before crawling. Google also explains that a file’s rules apply to the host, protocol, and port where that file is served. These statements describe Google’s crawler guidance, not a complete rulebook for every scraper. Google: robots.txt specifications.
Rank #4
Before scraping a particular site, separately review its terms, access controls, privacy implications, applicable law, and your intended use of the data. A robots.txt rule alone does not determine whether a particular scrape is permitted.
Troubleshoot common failures
- The selector times out: confirm the selector against the live DOM and check whether the content appears only after an interaction, a response, or scrolling. Wait for the meaningful condition rather than extending a fixed sleep without evidence.
- The page loads but extracted fields are null: inspect the item markup; the site may have changed its structure, or the field may be absent for some items. Make parsing tolerant of missing values and validate a sample.
- The browser shows fewer results than expected: determine whether the page paginates or loads more items on scroll. A visible first batch is not proof that the whole collection is present.
- The response wait never resolves: verify the page actually triggers the request, and refine the URL and status match using observed network activity. The example’s
/api/catalogpattern is illustrative, not a universal endpoint. - Navigation exceeds the timeout: distinguish a navigation failure from a page that loaded its initial document but is still fetching content. Use a navigation condition and a separate wait for the specific content needed.
- A managed crawl behaves differently from a direct session: check which Cloudflare Browser Run option you selected; Quick Actions, browser sessions, and the crawl endpoint serve distinct use cases.
Or skip the browser setup
If your goal is a screenshot or PDF of a dynamic page rather than structured field extraction, ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. Its clean-shot steps accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients.
For example, install Python’s requests package and run:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
See the ScreenshotNeo API documentation for authentication and request options. This captures a page image; it is not a substitute for extracting structured records from a site’s data request or DOM.
Best Value
ScreenshotNeo’s Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
Frequently Asked Questions
Can a scraper retrieve data from a page that requires JavaScript?
Yes. Reproduce the request that supplies the data when practical, or use browser automation to run the page when its JavaScript or interaction is necessary.
Does robots.txt grant permission to scrape a website?
No. It communicates crawler rules within its scope; review the site’s terms, access controls, privacy concerns, applicable law, and intended data use separately.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




