The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Puppeteer or Playwright when the page must execute JavaScript, maintain a session, or expose browser-only state. Use a scraping API when you would rather send a URL to a managed service than install, patch, scale, and monitor browsers yourself. A hosted-browser WebSocket connection preserves most of your existing script; a stateless REST endpoint is simpler for one-off rendering or extraction but gives you less control.
This guide shows both models, explains where each fits, and covers cleanup, memory, proxies, retries, and failure modes. It also separates browser automation from a screenshot-only API, because those are different jobs.
When browser rendering is necessary
A plain HTTP client can fetch server-rendered HTML quickly. It cannot, by itself, run the JavaScript that builds a product grid, waits for an API response, opens a menu, or applies client-side authentication. A browser is appropriate when the data appears only after scripts execute or when extraction depends on the same interactions a visitor performs.
- Use direct HTTP first for static HTML, feeds, documented JSON endpoints, or pages whose data is present in the initial response.
- Use Puppeteer or Playwright for JavaScript-rendered content, clicks, scrolling, authenticated sessions, screenshots, downloads, and browser network inspection.
- Use a task API when the required output is a rendered document, screenshot, PDF, or selector-based extraction and you do not need a long-lived browser object.
Rendering costs more resources than downloading HTML. Apify’s platform documentation states that Actors using Puppeteer or Playwright for real browser rendering require at least 1024 MB of memory. That is an Apify platform requirement, not a universal minimum for every machine or provider.
#1 Best Overall
Three ways to run a browser-backed scraper
| Approach | Your responsibility | Best fit | Main trade-off |
|---|---|---|---|
| Local automation | Install and launch the browser, manage versions, CPU, memory, concurrency, and cleanup. | Maximum control, private networks, custom workflows, and persistent sessions. | Infrastructure and browser maintenance remain yours. |
| Managed browser over WebSocket | Keep your Puppeteer or Playwright code; connect it to a provider’s browser. | Existing scripts that need hosted compute, geographic routing, or centralized operations. | Connection protocol, latency, session limits, data handling, and billing depend on the provider. |
| Stateless HTTP endpoint | Send a URL and extraction/rendering options; receive a response. | One-shot content, screenshots, PDFs, or selector extraction. | Less interactive control and provider-specific input/output limits. |
Browserless documents all three patterns: managed browser connections for existing scripts and REST endpoints for stateless tasks. Its REST catalogue includes smart scraping, rendered content, CSS-selector extraction, screenshots, PDFs, downloads, function execution, and unblocking. Choose an endpoint by output, not by the vague label “scraping API.”
Scrape a JavaScript page locally with Puppeteer
Puppeteer is a JavaScript library for controlling Chrome or Firefox through the DevTools Protocol or WebDriver BiDi. Its documentation says it runs headless by default. The following example launches a browser, waits for a product selector, extracts text, and always closes resources.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({headless: true});
try {
const page = await browser.newPage();
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 60000
});
await page.waitForSelector('.product-card', {timeout: 30000});
const products = await page.$$eval('.product-card', cards =>
cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null
}))
);
console.log(JSON.stringify(products, null, 2));
} finally {
await browser.close();
}
Use a selector wait for a known readiness condition instead of an arbitrary sleep. For pages that continue loading images or data, combine a selector wait with a bounded timeout. Puppeteer’s page APIs also support page content and screenshots, so the same session can extract and archive evidence.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Connect Puppeteer to a managed browser
For a hosted browser, install puppeteer-core and replace local launch with connect(). The provider supplies a WebSocket endpoint and authentication token.
import puppeteer from 'puppeteer-core';
const browser = await puppeteer.connect({
browserWSEndpoint: process.env.BROWSER_WS_ENDPOINT
});
try {
const page = await browser.newPage();
await page.goto('https://example.com/catalog', {waitUntil: 'networkidle2', timeout: 60000});
await page.waitForSelector('.product-card', {timeout: 30000});
console.log(await page.$$eval('.product-card', cards =>
cards.map(c => c.textContent.trim())
));
} finally {
await browser.close();
}
Closing the browser in finally matters. Browserless warns that an abandoned session can remain active until timeout and consume units; the exact billing behavior is provider-specific.
Scrape with Playwright
Playwright’s browser API documents Chromium, Firefox, and WebKit, with launch and session configuration. Its locator model and context isolation are useful when several tasks must run without sharing cookies.
import { chromium } from 'playwright';
const browser = await chromium.launch({headless: true});
try {
const context = await browser.newContext({
viewport: {width: 1440, height: 900}
});
const page = await context.newPage();
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 60000
});
await page.locator('.product-card').first().waitFor({timeout: 30000});
const products = await page.locator('.product-card').evaluateAll(cards =>
cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null
}))
);
console.log(products);
} finally {
await browser.close();
}
Connect Playwright over CDP
Browserless documents connectOverCDP() because its endpoint speaks the Chrome DevTools Protocol rather than Playwright’s own server protocol.
import { chromium } from 'playwright';
const browser = await chromium.connectOverCDP(process.env.BROWSER_CDP_ENDPOINT);
try {
const context = browser.contexts()[0] ?? await browser.newContext();
const page = await context.newPage();
await page.goto('https://example.com/catalog', {waitUntil: 'networkidle', timeout: 60000});
await page.locator('.product-card').first().waitFor({timeout: 30000});
console.log(await page.locator('.product-card').allTextContents());
} finally {
await browser.close();
}
Network control, sessions, and proxies
Playwright documents network monitoring and modification, including request and response handlers. Use it to observe the JSON call that populates a page, block analytics you do not need, or provide test fixtures. A browser-level route is not a substitute for permission to collect data.
Rank #3
Playwright also documents HTTP(S) and SOCKS v5 proxy configuration globally or per browser context. A proxy changes where traffic is routed; it does not make collection lawful, bypass a site’s terms, or guarantee access. Check the target site’s terms, robots guidance where applicable, authentication requirements, and applicable law.
Keep cookies and local storage isolated with a new context per account or job. Reuse a context only when continuity is intentional. Record the URL, status, timing, and extraction count so an empty result is distinguishable from a valid empty page.
Choosing a REST scraping endpoint
A REST request is attractive when the task has a clear input and output: URL in, HTML, structured fields, screenshot, or PDF out. Browserless describes separate paths for smart scrape, rendered content, CSS-selector scrape, and screenshots. Compare an endpoint on selector support, JavaScript execution, authentication inputs, output size, retry semantics, and handling of blocked or incomplete pages.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11- Define the output contract before choosing an endpoint: raw rendered HTML, selected fields, an image, or a PDF.
- Set an explicit timeout and retry only transient failures. Do not blindly retry a deterministic selector mismatch.
- Persist the response metadata and target URL. This lets you diagnose redirects, consent pages, and empty states.
- Rate-limit jobs and cap concurrency according to the provider’s documented limits and your target’s capacity.
Performance, reliability, and cost decisions
- Browser startup: launching a browser is slower and heavier than an HTTP request. Keep a controlled pool for sustained workloads, but close contexts and pages after each job.
- Waiting: prefer a selector, response, or network-idle condition tied to the page’s behavior. Always retain a maximum timeout.
- Concurrency: more pages increase memory and CPU pressure. Measure your workload rather than assuming a fixed number of tabs is safe.
- Retries: retry navigation timeouts and temporary server errors with backoff; capture a diagnostic screenshot or HTML when a parse fails.
- Sessions: persistent login and carts require a managed context or saved state. Stateless APIs generally start fresh unless they explicitly support cookies or session identifiers.
- Billing: hosted-browser units, session duration, REST calls, and bandwidth are provider-specific. The reviewed documentation does not establish a neutral price, success-rate, or speed comparison.
Common failures and fixes
“Selector not found”
The selector may be wrong, the page may have redirected to consent or login, or the data may be inside an iframe or shadow root. Save the final URL and HTML, verify the selector in browser developer tools, and wait for the API response or frame that creates the element.
Timeout during navigation
Large assets, a slow origin, or a blocked request can prevent the chosen load condition. Increase the bounded timeout only after identifying the slow stage; consider domcontentloaded followed by a specific readiness wait.
Empty results with HTTP 200
HTTP success does not prove that the intended state loaded. Check for bot checks, login forms, consent overlays, region-dependent content, and client-side errors. Log a screenshot and the rendered HTML.
Remote connection closes
Check the WebSocket or CDP URL, token, protocol, and provider session limits. Ensure your code closes contexts and browsers, and avoid holding an idle session longer than necessary.
Memory exhaustion
Reduce concurrency, block unnecessary resource types, reuse a browser while isolating contexts, and close pages promptly. On Apify, remember that its stated browser-Actor minimum is 1024 MB.
Best Value
Or skip the browser setup
If your deliverable is a clean screenshot rather than extracted records, ScreenshotNeo makes a single HTTP request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for all options. The service supports PNG, JPEG, WebP, and PDF; full-page capture with lazy-image loading; CSS-selector elements; dark mode; 12 device presets or custom viewports; retina scale; PDF paper, margins, orientation, and page ranges; custom CSS and JavaScript; pre-capture clicks; hidden selectors; selector, delay, or network-idle waits; ad, tracker, request, and resource blocking; custom headers, cookies, user agents, Authorization, timezone, and geolocation; transparent backgrounds; resizing; configurable-TTL caching; signed public-image links; asynchronous jobs with signed webhooks; bulk capture of up to 100 URLs per call; a usage API; and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Recommended Free Tools
A practical decision checklist
- Is the content already in server HTML or JSON? Use direct HTTP.
- Do you need clicks, login state, scrolling, downloads, or network interception? Use Puppeteer or Playwright.
- Do you have an existing script but want hosted compute? Use a managed WebSocket browser and keep explicit cleanup.
- Do you need one rendered artifact with little interaction? Choose a REST endpoint whose output matches the task.
- Will jobs run continuously? Budget memory, concurrency, session duration, retries, and provider billing before deployment.
- Are you capturing images rather than extracting records? Use a screenshot API such as ScreenshotNeo instead of maintaining a browser scraper.
Frequently Asked Questions
Can I use Playwright with Firefox or WebKit on a hosted browser?
Playwright documents Chromium, Firefox, and WebKit locally. A hosted connection exposes only the browser engines and protocol that that provider documents, so verify engine availability before designing the deployment.
Does a proxy make scraping permitted?
No. A proxy changes network routing only. Permission, terms, authentication, privacy obligations, and applicable law still govern collection.
Should every scraper use a browser?
No. Browser rendering is justified by client-side behavior or interaction requirements; static HTML and documented data endpoints are usually simpler and lighter.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

