Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use a browser-automation library such as Playwright to open the page, wait for the specific content you need, locate it with a user-facing locator, and read its text or attributes. A page reaching the load event is not proof that lazy or client-rendered data is ready. The reliable sequence is navigation, content-specific readiness, resilient locators, and extraction.
What you need before collecting data
- Node.js and a Playwright project.
- A target URL you are permitted to access and a clear definition of the fields to collect.
- A readiness signal for the data, such as a visible heading, a result count, or a loading indicator disappearing.
- A stable locator for each field. Prefer what a user can see or understand rather than a long chain of implementation-specific elements.
Install Playwright in a new project:
mkdir site-capture
cd site-capture
npm init -y
npm install playwright
npx playwright install
Playwright documentation describes locators as “the central piece of Playwright’s auto-waiting and retry-ability.” A locator is resolved when it is used, so it can continue to work when a page re-renders. See the Playwright locator guide.
A complete Playwright capture
The following script captures product cards after the page has rendered them. Replace the URL and the locator names with those from your permitted target.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
try {
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 30_000
});
// Wait for the data itself, not merely for navigation to finish.
const cards = page.getByRole('article');
await cards.first().waitFor({ state: 'visible', timeout: 20_000 });
const records = await cards.evaluateAll(elements => elements.map(card => ({
title: card.querySelector('[data-testid="product-title"]')?.textContent?.trim() || null,
price: card.querySelector('[data-testid="product-price"]')?.textContent?.trim() || null,
href: card.querySelector('a')?.href || null
})));
console.log(JSON.stringify(records, null, 2));
} finally {
await browser.close();
}
})();
Save it as capture.js and run node capture.js. The output is an array of records with a title, price, and absolute link. If the site uses a different accessible structure, choose a locator that matches the page’s actual semantics.
#1 Best Overall
Step 1: Navigate to the page
Use page.goto() with a timeout appropriate for the site. waitUntil: 'domcontentloaded' only waits for the initial document; it does not claim that API responses, images, or client-rendered components are complete. Playwright’s navigation guidance explains why lazy-loaded content can appear after navigation.
await page.goto('https://example.com/account/orders', {
waitUntil: 'domcontentloaded',
timeout: 30_000
});
For links that trigger a navigation, combine the action and navigation wait when needed:
await Promise.all([
page.waitForURL('**/results**'),
page.getByRole('button', { name: 'Search' }).click()
]);
Do not treat a fixed sleep as your primary readiness strategy. A delay can be too short on a slow run and unnecessarily long on a fast one.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsStep 2: Wait for the target state
Wait for the element or state that proves your data is present. Examples include a result heading, a table row, or the disappearance of a spinner.
Wait for a visible result
const heading = page.getByRole('heading', { name: /results/i });
await heading.waitFor({ state: 'visible' });
Wait for a loading indicator to disappear
await page.locator('[aria-label="Loading"]').waitFor({
state: 'hidden',
timeout: 20_000
});
Wait for a known count or value
const rows = page.getByRole('row');
await expect(rows).toHaveCount(11);
If you use assertions, import Playwright’s assertion helper:
const { chromium, expect } = require('playwright');
Choose a condition tied to the data contract. “Network idle” can be useful for pages that finish their requests, but an application may keep analytics or long-polling requests open; a visible, content-specific condition is usually clearer.
Step 3: Choose resilient locators
Playwright recommends user-facing attributes such as roles, text, labels, placeholders, alternative text, and titles. These communicate intent and are generally less coupled to internal DOM structure than generated class names or deeply nested CSS.
Common locator choices
| Need | Example | When to use |
|---|---|---|
| Button or link | page.getByRole('button', { name: 'Save' }) |
Interactive controls with accessible names |
| Form field | page.getByLabel('Email') |
Inputs associated with visible labels |
| Visible wording | page.getByText('In stock') |
Distinct text that users see |
| Placeholder | page.getByPlaceholder('Search products') |
Inputs whose placeholder is stable |
| Image | page.getByAltText('Company logo') |
Meaningful alternative text |
| Specific hook | page.locator('[data-testid="product-title"]') |
A deliberate test or data attribute exists |
CSS and XPath remain useful when no user-facing or dedicated hook exists. Keep them short and local. A selector such as div:nth-child(3) > div > span can break when an unrelated wrapper is inserted.
Disambiguate matches
const firstCard = page.getByRole('article').first();
const namedCard = page.getByRole('article').filter({
hasText: 'USB-C charger'
});
Make sure your locator identifies the intended element. A broad text locator may match navigation, hidden content, and the result you want at the same time.
Step 4: Extract text and attributes
One element
const title = await page.getByRole('heading', { name: /product/i }).innerText();
const canonical = await page.locator('link[rel="canonical"]').getAttribute('href');
Use innerText() when you want rendered, user-visible text. Use textContent() when hidden text and whitespace behavior are appropriate to your data contract. getAttribute() reads values such as href, src, datetime, or a custom data attribute.
One matched element with evaluate()
const sku = await page.getByRole('article').first().evaluate(card => ({
sku: card.getAttribute('data-sku'),
image: card.querySelector('img')?.currentSrc || null
}));
evaluate() runs a function in the page for the element matched by the locator. Keep the returned value serializable: strings, numbers, booleans, arrays, and plain objects.
Free tools Windows power users keep installed
One-click scans. No signup required.
A collection with evaluateAll()
const links = await page.getByRole('link').evaluateAll(anchors =>
anchors.map(a => ({
text: a.textContent.trim(),
href: a.href
}))
);
For a small number of elements, you can also use allTextContents() or allInnerTexts(). The Locator API documents these methods and the evaluation methods.
Rank #3
Capturing a dynamic list safely
locator.all() does not wait for matches. If the list is still changing, the returned set can be incomplete or inconsistent. First establish readiness, then collect.
const items = page.getByRole('listitem');
await page.locator('[data-testid="list-loading"]').waitFor({ state: 'hidden' });
await items.first().waitFor({ state: 'visible' });
const handles = await items.all();
const output = [];
for (const item of handles) {
output.push({
text: await item.innerText(),
href: await item.getAttribute('data-url')
});
}
If the application paginates or virtualizes rows, decide whether you need only the visible page or must click “Next” and repeat extraction. For infinite scrolling, scroll, wait for an explicit count or sentinel, and stop when the count no longer increases. Set a maximum page or item limit so a faulty page cannot create an unbounded job.
Attributes, normalization, and output
Normalize at the boundary rather than mixing page interaction with business logic:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →function clean(value) {
return value == null ? null : value.replace(/s+/g, ' ').trim();
}
const data = await page.getByRole('article').evaluateAll(cards =>
cards.map(card => ({
name: clean(card.querySelector('h2')?.textContent),
url: card.querySelector('a')?.href || null,
published: card.querySelector('time')?.getAttribute('datetime') || null
}))
);
In this example, define clean outside the browser callback if your runtime cannot serialize a referenced function into the page. A safer version passes only DOM operations to evaluateAll(), then normalizes the returned strings in Node.js:
const raw = await page.getByRole('article').evaluateAll(cards =>
cards.map(card => ({
name: card.querySelector('h2')?.textContent || null,
url: card.querySelector('a')?.href || null
}))
);
const data = raw.map(row => ({
...row,
name: row.name?.replace(/s+/g, ' ').trim() || null
}));
Write JSON with require('fs').writeFileSync('data.json', JSON.stringify(data, null, 2)), or send records to your database after validating required fields.
Authentication, privacy, and responsible operation
- Use an account and session you are authorized to automate. Do not bypass access controls, CAPTCHAs, or bot protections.
- Keep credentials out of source control. Playwright can load secrets from environment variables.
- Collect only the fields you need, protect personal data, and follow the site’s terms and applicable law.
- Throttle repeated requests and cache results when freshness permits. A browser is resource-intensive; close each browser and context in a
finallyblock.
const context = await browser.newContext({
extraHTTPHeaders: { Authorization: `Bearer ${process.env.API_TOKEN}` }
});
Troubleshooting common failures
“Timeout exceeded” while waiting
The locator may be wrong, the page may require an interaction, or the site may be slow. Inspect the page with await page.screenshot({ path: 'debug.png', fullPage: true }), log await page.url(), and verify the locator in headed mode. Increase the timeout only after confirming the condition is correct.
The script sees zero list items
You likely collected before the client-rendered list appeared, or the list is inside an iframe. Wait for a list-specific signal. For a frame, obtain it with page.frameLocator('iframe[title="Results"]').getByRole('listitem').
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Text is empty or different from the screen
Check whether the value is in an attribute, shadow DOM, or a child that has not rendered. Try getAttribute(), a more specific locator, or a page-state wait. Remember that textContent() and innerText() have different visibility and whitespace behavior.
Results change between runs
The page may personalize, paginate, or update continuously. Fix the input state, wait for a stable condition, capture a timestamp, and record the URL and extraction count. Avoid relying on positional selectors.
Browser launch fails in deployment
Install the browser binaries during the build, confirm the runtime has the required dependencies, and use a supported headless configuration. Log the Playwright version and browser launch error rather than silently returning an empty dataset.
Performance and reliability checklist
- Reuse one browser process for multiple pages, but isolate cookies and storage in separate contexts when jobs must not share sessions.
- Block unnecessary resources only when doing so cannot remove data you need.
- Use bounded timeouts, retries for transient navigation failures, and idempotent output writes.
- Record the target URL, capture time, item count, and an error reason for every job.
- Validate a sample of extracted records against the page so a selector change does not produce plausible but wrong data.
Or skip the browser setup
If your goal is a screenshot or PDF rather than structured DOM records, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL in one GET request and can return PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
Recommended Free Tools
For an API capture, see the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The equivalent Python request is:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Options include full-page capture with lazy images loaded, CSS-selector element capture, device presets, custom viewport and retina scale, PDF paper and page controls, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. It supports parameter names used by other screenshot APIs to ease migration.
Best Value
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000 per month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing provides two months free, and every feature is available on every plan. Sign up free for 1,000 screenshots a month with no card.
When browser extraction is the right tool
Use Playwright when you need structured fields from rendered interfaces, interactions such as search or pagination, or data that exists only after JavaScript runs. Use a screenshot API when the deliverable is a visual record or PDF and you do not need to parse individual nodes. For either approach, define readiness explicitly and verify the output instead of assuming navigation success means data success.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Frequently Asked Questions
Can Playwright capture data from a page that requires scrolling?
Yes. Scroll or trigger the page’s pagination, wait for a content-specific signal after each action, then extract the newly available elements. Set a maximum number of iterations.
Should I use CSS selectors or XPath for every field?
No. Start with role, label, text, placeholder, alt text, or title locators. Use CSS or XPath when the page offers no reliable user-facing or dedicated attribute, and keep the selector as local as possible.
What is the difference between a screenshot and data extraction?
Extraction returns structured values such as text, links, and attributes. A screenshot or PDF preserves rendered appearance. Choose based on the output your application must consume.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

