Recommended Free Tools
Use a real browser to extract JavaScript-rendered data: open the page with Playwright, wait for the element that proves the data is ready, locate records with resilient selectors, evaluate only the fields you need, and validate the result before saving it. This approach captures the DOM a visitor sees instead of the incomplete HTML often returned by a basic HTTP request.
When browser automation is the right extraction method
Start by checking for an official API, export, or structured feed. A supported interface is usually simpler and less fragile than driving a visible page. Use browser automation when the values appear only after JavaScript runs, require scrolling or interaction, or are assembled from client-side requests that you cannot reasonably consume directly.
The examples below use Playwright with Node.js. The same workflow applies to Python and other Playwright bindings: navigate, wait for meaningful state, select, extract, validate, and handle failure explicitly. Check the target site’s terms and applicable rules for your project. Robots directives are crawler-facing guidance for cooperative crawlers; they do not by themselves settle permission or legal questions.
Install Playwright and create a first extractor
- Create a project:
mkdir page-extract && cd page-extract && npm init -y. - Install the library and browser:
npm install playwright, thennpx playwright install chromium. - Save this script as
extract.js:
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
viewport: { width: 1440, height: 900 }
});
try {
await page.goto('https://example.com/products', {
waitUntil: 'domcontentloaded',
timeout: 45_000
});
const cards = page.locator('[data-testid="product-card"]');
await cards.first().waitFor({ state: 'visible', timeout: 15_000 });
const products = await cards.evaluateAll(nodes => nodes.map(node => ({
name: node.querySelector('[data-testid="product-name"]')?.textContent?.trim() || null,
price: node.querySelector('[data-testid="product-price"]')?.textContent?.trim() || null,
url: node.querySelector('a')?.href || null
})));
if (products.length === 0) {
throw new Error('No product cards matched; refusing to save an empty dataset');
}
for (const product of products) {
if (!product.name || !product.url) {
throw new Error(`Incomplete record: ${JSON.stringify(product)}`);
}
}
console.log(JSON.stringify(products, null, 2));
} finally {
await browser.close();
}
})();
Replace the URL and selectors with the page’s actual structure. The data-testid attributes above are illustrative. Inspect a representative record in browser developer tools and choose selectors that identify the intended content, not nearby navigation or advertising.
#1 Best Overall
Wait for rendered data, not just navigation
domcontentloaded means the initial document was parsed; it does not mean a client-rendered list has arrived. Wait for a meaningful list, card, heading, or application state:
await page.goto(target, { waitUntil: 'domcontentloaded' });
await page.getByRole('heading', { name: 'Search results' }).waitFor();
await page.locator('[data-testid="result-row"]').first().waitFor({ state: 'visible' });
Prefer an element-based wait over an arbitrary sleep. If the application exposes a stable signal such as a loading indicator disappearing, wait for that state as well. For content that updates repeatedly, wait until the expected count is reached or poll until the count remains unchanged for a short interval. A multiple-element query such as locator.all() does not itself wait for a dynamic list to finish; establish readiness first.
Waiting for interaction-driven content
Click the control that reveals the data, then wait for the resulting region:
await page.getByRole('button', { name: 'Load more' }).click();
await page.locator('[data-testid="result-row"]').nth(19).waitFor();
For infinite scrolling, scroll in bounded steps and stop when no new records appear. Record the final count so a site change cannot silently truncate your run.
Choose selectors that survive redesigns
Playwright recommends user-facing locators because they describe what a visitor or assistive technology can identify:
getByRole('button', { name: 'Export' })for controls.getByLabel('Email')for form fields.getByText('Acme Corporation')when the visible text is the data anchor.getByTestId('result-row')when the site deliberately provides a test contract.
Use concise CSS for batch extraction when no semantic locator exists. Avoid long chains such as div:nth-child(3) > div > span; they encode implementation details and tend to break after a redesign. XPath can express an awkward relationship, but a long structure-dependent path is difficult to maintain. Locators are strict for operations that imply one target: if two buttons match, Playwright can raise an error. Narrow the region or accessible name rather than hiding ambiguity with first() or nth().
Extract text, links, attributes, and structured values
Locator evaluation
Evaluate a focused function in the page context when you need to transform matched nodes:
const rows = page.locator('table tbody tr');
await rows.first().waitFor();
const data = await rows.evaluateAll(trs => trs.map(tr => {
const cells = [...tr.querySelectorAll('td')];
return {
title: cells[0]?.textContent?.trim() ?? null,
status: cells[1]?.textContent?.trim() ?? null,
href: tr.querySelector('a')?.getAttribute('href') ?? null
};
}));
Return plain serializable objects rather than DOM nodes. Normalize whitespace, parse numbers deliberately, and preserve the original link when it is useful for auditing.
CSS selection in the browser DOM
const links = await page.evaluate(() => [...document.querySelectorAll('article a')]
.map(a => ({ text: a.textContent.trim(), href: a.href })));
MDN’s querySelectorAll() returns a static NodeList in document order. It will not update after a click, pagination request, or virtual-list render; run the query again after each relevant page change. Invalid CSS syntax throws an error, and unusual IDs or class names may need escaping.
Validate before you trust the dataset
- Assert an expected range or minimum count.
- Require key fields and reject malformed URLs.
- Check duplicates using a stable ID or canonical URL.
- Log a few representative records and the page URL.
- Save an HTML snapshot or screenshot when diagnosing a mismatch.
- Treat zero matches as a failure signal, not a successful empty export.
Pagination, lazy loading, consent dialogs, and interaction gates are site-specific. Verify that you collected every page or cursor, and rerun selectors after content changes.
Rank #3
Common failures and precise fixes
The list is empty
Cause: the query ran before rendering, the selector drifted, or a consent dialog blocked the page. Fix: inspect the live DOM, wait for a specific record, handle the dialog if permitted, and log the final URL and page title.
Only the first page was captured
Cause: pagination or infinite scrolling was not implemented. Fix: loop through the Next control or documented cursor, wait for the old page to become stale or the count to increase, and stop on a disabled control or repeated cursor.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Results are duplicated or stale
Cause: a static querySelectorAll() result was reused after an update. Fix: query again after each interaction and deduplicate by a stable key.
A locator matches several elements
Cause: the selector is not specific enough. Fix: scope it to the relevant card, region, role, or accessible name. Do not use positional methods merely to conceal ambiguity.
Navigation times out
Cause: slow resources, a blocked request, or a bot challenge. Fix: set a realistic timeout, capture a diagnostic screenshot, inspect response status and console errors, and decide whether the site permits automated access. A longer timeout cannot solve a challenge page.
Fields are present visually but missing in output
Cause: the value may be in an attribute, shadow DOM, an iframe, or a later render. Fix: use getAttribute(), inspect frames with page.frames(), wait for the field itself, and confirm whether the component exposes accessible text.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteReliability, performance, and responsible operation
Reuse one browser process and create isolated pages or contexts for batches. Keep concurrency bounded so the target and your machine are not overwhelmed. Block unnecessary images or analytics only when doing so cannot alter the data you need. Cache results with a timestamp, retry transient navigation failures with backoff, and record status, duration, selector version, and error details. A failed or partial run should be visible in logs and should not overwrite the last known-good export.
Virtualized lists may contain only visible rows in the DOM; scroll and collect incrementally. Shadow DOM and cross-origin iframes require component- or frame-specific handling. Authentication, geolocation, cookies, and custom headers must be supplied only when you are authorized to use them. Respect rate limits and revisit selectors after the site changes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo provides a website screenshot API when your goal is a visual capture rather than a custom DOM dataset. One GET request returns PNG, JPEG, WebP, or PDF. For example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options and response headers. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status. Its MCP server lets Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe Free plan includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
Best Value
Python and Node.js alternatives
Python HTTP capture with ScreenshotNeo
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js HTTP capture with ScreenshotNeo
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));
These calls produce an image or PDF, not the structured records produced by Playwright. Choose the API when a clean visual artifact is sufficient; keep browser extraction for field-level data, pagination, and custom transformations.
Frequently Asked Questions
Should I use an API instead of browser automation?
Yes, when the site offers an authorized API, export, or feed containing the fields you need. Use browser automation when the rendered interface is the available source.
Why does a fixed sleep often fail?
A sleep guesses timing. Network speed and rendering vary, so wait for the specific element or state that proves the data is ready.
Free tools Windows power users keep installed
One-click scans. No signup required.
Can robots.txt authorize my scraper?
No. Robots directives guide cooperative crawlers, but they are not a complete permission or legal analysis for a particular project.
What should I store for reproducibility?
Store the URL, retrieval time, selector version, record count, errors, and enough raw context—such as a snapshot or screenshot—to explain unexpected output.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

